Edition 2026.07b

    State of the field

    A dated snapshot of what this discipline can currently claim, sorted by how much confidence the evidence supports. Sections are updated in place; the edition number changes when they do.

    Nothing here is presented as more settled than it is. The contested and open sections exist because a reference that only publishes its confident claims is not a reference.

    Currency
    Last reviewed end-to-end on . Claims in “newly observable” carry the shortest half-life and should be treated as expired if this date is more than a quarter old.

    Edition history

    • 2026.07b · July 29, 2026Added a settled claim on speed and rigour (Eisenhardt), a contested entry on whether synthesis tooling improves or launders qualitative evidence, and an open question on how machine-generated signals should be weighted against human-observed ones. Published the PROBE Spec, Calibration Card, and DRIFT Check.
    • 2026.07 · July 26, 2026First edition. Four sections, sourced claims, open questions published unresolved.

    What looks settled

    Positions that have survived enough scrutiny that arguing against them requires new evidence rather than a new opinion.

    Expert intuition is reliable only under specific, nameable conditions.

    Kahneman and Klein's joint paper is the closest thing this field has to a settled result: intuition can be trusted where the environment is regular enough to contain learnable patterns and where the practitioner has had the opportunity to learn them through feedback. Outside those conditions, confidence and accuracy come apart. This is the ceiling on every 'trust your gut' argument, and the reason vibe science insists on writing the gut down before testing it.

    Conditions for Intuitive Expertise (2009)

    Different problem domains require structurally different responses.

    Cynefin's distinction between complicated (analysable by experts) and complex (only knowable through probing) has held up for nearly two decades of practitioner use. Most failed strategy work is a category error: applying analysis to a domain that only responds to probes.

    A Leader's Framework for Decision Making (HBR, 2007)

    Forecasting accuracy is trainable, and the training is mostly bookkeeping.

    The Good Judgment Project's central finding — that ordinary people who record predictions, score them, and update methodically outperform credentialled analysts — is the empirical backbone of the argument that dated, written intuitions compound into judgement while undated ones do not.

    Good Judgment Project

    Teams without psychological safety cannot surface weak signals at all.

    Edmondson's research is unusually robust here: in environments where saying 'something feels wrong' without data carries a cost, the signals do not get louder, they get private. The organisational prerequisite is not a process — it is permission.

    Amy Edmondson

    Deciding faster does not require deciding with less.

    Eisenhardt's field research in high-velocity environments found that the fastest strategic decision makers used more real-time information and considered more alternatives than slower ones — the delay in slow teams came from deferring the decision, not from the analysis. This is the closest thing to evidence that speed and rigour are not the trade-off they are usually assumed to be, and it is why our frameworks are measured in minutes rather than workshops.

    Making Fast Strategic Decisions in High-Velocity Environments (1989)

    What is genuinely contested

    Live disagreements where confident people are on both sides and the evidence does not yet decide it. Listed because pretending otherwise is how a reference page goes stale.

    Whether AI systems can meaningfully assist in the framing of a problem, or only in its execution.

    The optimistic position is that models are already useful adversaries — good at generating the null-case you did not want to consider. The sceptical position is that they regress toward the most common framing in the training data, which is precisely the wrong instinct at the frontier. Both camps have practitioners with real track records. Our working stance is narrow: use models to attack a hypothesis you wrote, not to write it.

    Whether agentic systems reduce or increase the cost of being wrong.

    Autonomy compresses the loop between idea and action, which makes experiments cheaper. It also introduces silent partial failures, which makes wrongness harder to detect. As of mid-2026 the deployment data is thin and heavily vendor-sourced; treat confident claims in either direction as marketing until independent evaluations exist.

    Whether machine synthesis of qualitative material improves evidence or launders it.

    One camp holds that clustering hundreds of conversations surfaces patterns no tired human reader would catch, and that the alternative was never careful reading — it was reading twenty and generalising. The other holds that fluent output suppresses the questions that used to get asked of a summary, because it arrives already organised and already confident. Our position is procedural rather than partisan: use the synthesis, then run a trace check on it before it counts as evidence.

    The DRIFT check

    Whether benchmark performance transfers to domain work at all.

    The gap between leaderboard scores and observed usefulness on a specific team's real tasks has widened rather than closed. The practical consequence is that a private, hand-labelled evaluation set of fifty to a hundred examples remains more informative than any public benchmark — a position that is now common among practitioners but still argued about.

    Evaluation checklist

    Whether 'vibe' can survive as a serious term.

    The word carries a decade of connotation that works against the rigour the practice requires. The counter-argument is that memorable names travel and academic ones do not — data science won its name the same way. This site is running that argument as a public experiment rather than asserting the answer.

    The live naming experiment

    What is newly observable in 2026

    Changes recent enough that most existing playbooks predate them. These carry short half-lives by construction.

    Implementation cost has fallen far enough that it is no longer the binding constraint.

    When a working surface can be produced in an afternoon, the scarce resource stops being the ability to build and becomes the judgement to know what is worth building. This is the single largest structural change behind the case for a validation discipline, and it is why the throwaway probe has gone from a luxury to the default first move.

    The 2026 shift

    Tool interoperability has become an assumption rather than a project.

    Standardised context and tool-calling protocols mean the integration work that used to justify a quarter now justifies an afternoon. The strategic consequence: integration is no longer a moat, and any thesis that depended on it being one has expired.

    'Add an agent' has become the reflexive answer, exactly as 'add a chatbot' was in 2023.

    The useful discriminator is whether the problem is genuinely agent-shaped: multi-step work whose steps are not knowable up front, where the user's cost of specifying exceeds the model's cost of choosing, and where the environment tolerates silent partial failure. Most problems fail the third test.

    Field note: not every problem is agent-shaped

    The cost of collecting signal has fallen faster than the discipline to interpret it.

    Synthesis of interviews, tickets, and sessions is now cheap enough to run continuously. The failure mode has shifted accordingly: from having too little qualitative data to having a clustered summary nobody has traced back to a single verbatim sentence. Provenance is the control.

    The SPINE record

    Open questions we do not have an answer to

    Published rather than resolved. If you have evidence bearing on any of these, it is more valuable than agreement.

    What is the right base rate for killing bets?

    A team that kills nothing is not experimenting. A team that kills everything is not committing. There is no defensible published number for what a healthy kill rate looks like, by team size or domain, and we are not going to invent one.

    The Kill Ledger

    Do signal half-lives actually behave the way practitioners assume?

    The intuition that technical claims decay in weeks while claims about human motivation decay in years is widely shared and, as far as we can find, unmeasured at team scale. Our own estimates are published as a calibration prior and explicitly flagged as draft.

    Signal Half-Life

    Does structured signal logging improve decisions, or only improve the record of them?

    The honest answer is that we do not know. The record has independent value — it enables calibration that is otherwise impossible — but the claim that logging improves the decision itself is unproven, and we would rather say so than assert it.

    How should a machine-generated signal be weighted against a human-observed one?

    A model can now surface a pattern across ninety conversations that no person read end to end. A practitioner can notice one sentence in one call that reframes the quarter. There is no principled way we know of to weigh those against each other when they conflict, and the informal practice — trusting whichever agrees with the current plan — is clearly wrong. We are recording this as unresolved rather than proposing a weighting we cannot defend.