Frameworks

    Eight named models for working in the period before metrics exist. Each is written to be quoted, attributed, and argued with — one-paragraph definition, steps, a worked example, the prior art it builds on, and the conditions under which it breaks.

    These are proposals from this site, not established standards. They are versioned so that a citation made today still points at what was actually said.

    Framework hygiene
    Last reviewed . Frameworks marked draft contain estimates that have not been measured, and say so at the point where the estimate appears.

    The VSV Loop

    v1.0 · stable

    Vibe → Signal → Validation

    The VSV Loop is a three-phase cycle for working in the period before metrics exist: a Vibe is a named, dated intuition treated as the first draft of a hypothesis; a Signal is the earliest observable evidence that the hypothesis is or is not tracking reality; Validation is the smallest test whose result would change the next decision. The loop is complete only when the result is written down — including when the result is that the idea was killed.

    Use when

    Any time you are acting on a read of a situation that your instruments cannot yet confirm or deny — a new market, a new behaviour, a new technology, a team dynamic nobody has quantified.

    1. V

      Vibe — name and date the intuition

      Write the observation, the observer, the date, and the context in one sentence. Unnamed intuitions cannot be scored later, which means they can never turn into calibrated judgement.

    2. S

      Signal — decide what would show up first

      Before looking for evidence, state what the earliest observable trace would be if the vibe were true, and what would be observable if it were false. Both, in advance. The second half is what stops confirmation bias.

    3. V

      Validation — run the smallest decisive test

      Smallest means: the cheapest test whose outcome changes your next action. If no outcome would change your action, you are seeking permission, not validating.

    Worked on this site's own live experiment

    • Vibe: there is an under-named role — the person doing structured inquiry before metrics exist — and a memorable label for it will travel further than an academic one.
    • Signal: unsolicited third-party uses of the phrase; people describing themselves in the term's language without being prompted.
    • Validation: at least one unaffiliated use of the phrase in a job description, bio, or public talk within 90 days. Kill criterion published alongside it, before the result was known.
    • Status: running in public at /experiments. The result is not known at the time of writing, which is the point.

    Built on

    Where it breaks

    VSV is a poor fit for the obvious and complicated domains, where best practice or expert analysis already answers the question faster. Running a loop where a known answer exists is procedural theatre.

    Cite thisVibe Scientist. “The VSV Loop (Vibe → Signal → Validation)”, v1.0. https://www.vibescientist.com/frameworks#vsv-loop. Retrieved July 29, 2026.

    The SPINE Record

    v1.0 · stable

    Signal · Provenance · Interpretation · Null-case · Expiry

    A SPINE record is a five-field note taken at the moment a weak signal is observed, containing the Signal (what was observed, verbatim), its Provenance (who observed it, where, when, and through what channel), the Interpretation (what you currently think it means, marked as interpretation), the Null-case (the mundane explanation that would make this nothing), and an Expiry (the date after which this record must be re-tested or discarded). A signal without all five fields is an anecdote with a timestamp.

    Use when

    Every time something makes you think 'huh' — a customer sentence, a support pattern, a competitor move, a metric that moved for no stated reason. Cost: under two minutes per record.

    1. S

      Signal — record the observation verbatim

      Quote it. Do not summarise it. Summaries encode your interpretation into what should be the raw data, and you will not be able to separate them again in six weeks.

    2. P

      Provenance — who, where, when, through what channel

      Provenance is what makes a signal auditable. It also exposes sampling: four signals from the same channel are one signal about that channel.

    3. I

      Interpretation — labelled as interpretation

      Write what you think it means, in a separate field, explicitly marked. The separation is the entire point: interpretations get revised, observations do not.

    4. N

      Null-case — the boring explanation

      State the most plausible way this is nothing: coincidence, a bad week, one loud user, a pricing page bug. If you cannot write a credible null-case, you are not looking hard enough for one.

    5. E

      Expiry — the re-test date

      Every record gets a date after which it is either re-confirmed or retired. Unexpired signals are how a team's model of the world quietly becomes a decade old.

    What a filled record looks like

    • Signal: 'We built the thing in two days. We just didn't know if anyone wanted it.' — said unprompted in a customer call.
    • Provenance: observed by the founder, 40-minute call, mid-market prospect, recorded with consent.
    • Interpretation: build cost has collapsed far enough that validation, not implementation, is now the binding constraint for this segment.
    • Null-case: one articulate person with an unusual stack, repeating a phrase they read online this month.
    • Expiry: re-test in 60 days — retire unless three unrelated sources produce the same shape of statement.

    Built on

    Where it breaks

    SPINE creates volume. A hundred records nobody reads is worse than ten records reviewed weekly, because it produces the feeling of rigour without the substance. Pair it with a standing review or do not adopt it.

    Cite thisVibe Scientist. “The SPINE Record (Signal · Provenance · Interpretation · Null-case · Expiry)”, v1.0. https://www.vibescientist.com/frameworks#spine. Retrieved July 29, 2026.

    The TRACE Filter

    v1.0 · stable

    Time-cap · Reversibility · Assumption · Cost-of-wrong · Evidence-needed

    TRACE is a five-question filter run before committing to any bet, answered in writing in under fifteen minutes: what is the Time-cap on this attempt; how Reversible is it and at what cost; what single Assumption must hold for it to work; what is the Cost-of-wrong across money, time, reputation, and morale; and what Evidence would be sufficient to proceed. A bet that cannot answer all five in fifteen minutes is not ready to be made — or is not actually a bet, just a habit.

    Use when

    Before any commitment that would be awkward to unwind: a launch, a hire, a rebuild, a partnership, a public position.

    1. T

      Time-cap

      The date the attempt ends whether or not it has succeeded. Without it, the bet has no failure state and will consume resource indefinitely by default.

    2. R

      Reversibility

      One-way door or two-way door, and the concrete cost of walking back. Most decisions treated as irreversible are reversible at modest cost; the bias runs strongly in one direction.

    3. A

      Assumption

      The single load-bearing belief. Not a list — the one that, if false, makes everything else irrelevant. Naming one forces prioritisation that a list avoids.

    4. C

      Cost-of-wrong

      Four numbers or four sentences: money, time, reputation, morale. Morale is the one teams skip and the one that compounds.

    5. E

      Evidence-needed

      What would be enough to proceed — specified before you go looking, so the threshold cannot drift down to meet whatever you happen to find.

    Applied to a common 2026 decision: 'should we add an agent?'

    • Time-cap: two weeks to a working probe, hard stop.
    • Reversibility: high for a probe, low for a rebuilt core workflow — so probe first, and do not let the probe quietly become the workflow.
    • Assumption: users cannot economically specify the steps themselves. If they can, this is a form, not an agent.
    • Cost-of-wrong: latency, per-action cost, and a new class of silent partial failures that erodes trust faster than it can be rebuilt.
    • Evidence-needed: on a fixed set of real tasks, autonomous completion beats the guided path on both success rate and time-to-outcome.

    Built on

    Where it breaks

    TRACE is a filter, not an analysis. It is designed to catch unexamined commitments cheaply. For genuinely irreversible, high-cost decisions it should escalate to a pre-mortem rather than substitute for one.

    Cite thisVibe Scientist. “The TRACE Filter (Time-cap · Reversibility · Assumption · Cost-of-wrong · Evidence-needed)”, v1.0. https://www.vibescientist.com/frameworks#trace. Retrieved July 29, 2026.

    The Kill Ledger

    v1.0 · stable

    A Kill Ledger is a permanent, dated register of every idea a team deliberately stopped, recording what was believed, what was observed, what criterion triggered the stop, and what was learned. It exists because killing an idea produces information that is otherwise discarded entirely: without a ledger, a team's abandoned bets leave no trace, cannot be revisited when conditions change, and cannot be used to calibrate whether that team kills too early, too late, or for the wrong reasons.

    Use when

    Standing artefact. One entry per killed bet, written the week the decision is made — not at quarter-end, when the reasoning has already been smoothed into a narrative.

    1. 01

      Record what was believed

      The original hypothesis in its original words. Not the revised, more defensible version.

    2. 02

      Record the triggering observation

      What was actually seen, and on what date. Distinguish 'the criterion was met' from 'we lost interest' — both are valid, only one is evidence.

    3. 03

      Record the criterion

      The kill criterion as written before the result was known. If none was written, note that too; it is the single most useful pattern to notice about yourself.

    4. 04

      Record the reversal condition

      What would have to change for this to be worth reopening. Many killed ideas are correct ideas at the wrong time.

    5. 05

      Review the ledger quarterly

      Read for pattern, not for individual verdicts: are the same failure modes recurring? Are kills happening after the money was spent? Is anything reopenable now?

    What the ledger tells you that nothing else does

    • A team with no kills is not disciplined — it is either not experimenting or not finishing.
    • A team whose kills all cluster after launch is deciding too late; the criteria are being written after the spend.
    • A team whose kills all cite 'deprioritised' rather than an observation is not running experiments; it is running a backlog.
    • A reopened entry, with the reversal condition met, is the highest-value item in the ledger — an idea that was right and early, recovered on time.

    Built on

    • Amy Edmondson — intelligent failureOn distinguishing failures that produce information from failures that produce only cost. The ledger is the record that makes the distinction visible.

    Where it breaks

    A ledger used for accountability rather than learning will be gamed within a quarter. If entries become performance-review inputs, people stop killing things, which is the exact opposite of the intent.

    Cite thisVibe Scientist. “The Kill Ledger”, v1.0. https://www.vibescientist.com/frameworks#kill-ledger. Retrieved July 29, 2026.

    Signal Half-Life

    v1.0 · draft

    Signal half-life is the period after which a signal's evidential value has decayed by roughly half — the point at which a conclusion drawn from it should be re-tested rather than inherited. It is an estimate, not a measurement: the practice is to assign every signal an explicit expiry at the moment it is recorded, so that beliefs held past their evidence become visible as expired records rather than invisible as institutional common sense.

    Use when

    At recording time for any signal, and at review time for any belief that is being used to justify a decision but cannot be traced to a recent observation.

    1. 01

      Classify the signal's rate of change

      Fast-moving domains (model capability, tooling, pricing norms, competitor behaviour) decay in weeks. Slow-moving ones (human motivation, organisational friction, category fundamentals) decay in years.

    2. 02

      Assign an explicit expiry date

      Write the date, not a duration. Durations are never computed; dates trigger.

    3. 03

      At expiry, re-test or retire

      Re-testing is usually cheap — one call, one query, one search. Retiring is free. Inheriting is the expensive option, and it is the default.

    4. 04

      Audit what your team believes but cannot date

      Ask, in a review: when did we last observe this? A room that cannot answer is running on an undated model of a changed world.

    Rough half-lives, offered as a starting calibration

    • Frontier model capability and cost: weeks. Anything you concluded about what models cannot do a year ago should be assumed expired.
    • Tooling and integration norms: months. Protocol and ecosystem conventions have been moving faster than documentation since 2025.
    • Competitor positioning: one to two quarters.
    • Why a specific user behaves the way they do: years, sometimes decades. Human motivation is the slowest-decaying signal you can collect, which is why interviews still outperform dashboards at the frontier.
    • These are estimates for calibration, not measured values. Treat them as a prior to be corrected by your own domain.

    Built on

    Where it breaks

    Marked draft deliberately: the half-life estimates above are calibration priors, not measured quantities. Anyone citing this should cite the practice — assign an expiry to every signal — rather than the specific durations.

    Cite thisVibe Scientist. “Signal Half-Life”, v1.0. https://www.vibescientist.com/frameworks#signal-half-life. Retrieved July 29, 2026.

    The PROBE Spec

    v1.0 · stable

    Premise · Reach · Observable · Budget · Exit

    A PROBE spec is the five-line brief written before building anything intended to generate evidence rather than value: the Premise being tested in one falsifiable sentence, the Reach (who must encounter it, and how many, for the result to mean anything), the Observable (the specific behaviour that counts as a result, decided in advance), the Budget (a hard ceiling in hours and money, not an estimate), and the Exit (what gets deleted, and when, regardless of outcome). A probe without an exit is a product nobody agreed to maintain.

    Use when

    Whenever the cheapest way to learn something is to put an artefact in front of real people — a landing page, a fake door, a manual concierge run, a throwaway agent. Especially when building is cheap enough that the temptation is to skip the brief entirely.

    1. P

      Premise — one falsifiable sentence

      Not 'test demand for X.' Rather: 'people in this segment will do Y when shown Z.' If you cannot phrase it so that a result could embarrass you, it is not a premise.

    2. R

      Reach — who, and how many

      State the population and the minimum count before you start. Small numbers are fine at the frontier; unstated numbers are not, because they let any result be read as encouraging.

    3. O

      Observable — the behaviour that counts

      Choose an action, not a sentiment. Enthusiasm in a call is not an observable. Booking time, paying, returning, or forwarding is.

    4. B

      Budget — a ceiling, not an estimate

      Hours and money, written down and agreed. Estimates expand to fit interest in the idea; ceilings do not. Crossing the ceiling ends the probe whether or not the answer arrived.

    5. E

      Exit — what gets deleted, and when

      Decide the teardown date at the start. Probes that quietly survive become infrastructure with no owner, and the next probe gets built on top of them.

    How the spec constrains a cheap build

    • Premise: operators who already keep decision notes will adopt a structured signal format if it costs under two minutes per entry.
    • Reach: people who have written publicly about decision journals; minimum twelve, recruited individually rather than broadcast.
    • Observable: a second entry logged without prompting, within fourteen days of the first.
    • Budget: sixteen hours of build, no paid distribution.
    • Exit: the artefact comes down on day thirty, result or no result, and the spec plus outcome goes in the kill ledger either way.

    Built on

    • Dave Snowden — safe-to-fail probesThe idea that experiments in complex domains must be bounded, parallel, and recoverable. PROBE adds the written pre-commitment and the teardown date.
    • Gary Klein — pre-mortemThe discipline of writing failure conditions before the work starts, rather than reconstructing them afterwards.

    Where it breaks

    PROBE is wrong for work whose value is the artefact itself. Applying an exit date to something customers depend on is not rigour, it is churn. It also underperforms where the observable behaviour has a long natural lag — enterprise procurement, clinical adoption — because the budget ceiling will fire before reality answers.

    Cite thisVibe Scientist. “The PROBE Spec (Premise · Reach · Observable · Budget · Exit)”, v1.0. https://www.vibescientist.com/frameworks#probe. Retrieved July 29, 2026.

    The Calibration Card

    v1.0 · stable

    A Calibration Card is a four-line record filed at the moment a judgement call is made: the claim in resolvable form, a probability expressed as a number, the date it resolves, and the single observation that would most change the number. Cards are scored on resolution and read in batches, never individually — the purpose is not to be right about any one call, but to learn the shape of your own error. A team that records numbers becomes calibrated within a few dozen resolved cards; a team that records adjectives never does.

    Use when

    Before any decision where you will later be tempted to say 'I always thought that would happen.' Particularly valuable for judgements made without data, which are exactly the ones memory rewrites most freely.

    1. 01

      State the claim so it can resolve

      'This will go well' cannot resolve. 'Three of the five pilots renew by 31 March' can. Resolvability is the whole cost of the method, and it is paid up front.

    2. 02

      Write a number, not an adjective

      Seventy per cent, not 'fairly confident.' Adjectives are unfalsifiable by design; that is why they are comfortable. Numbers feel arrogant for about three cards and then feel ordinary.

    3. 03

      Set the resolution date at the same moment

      Undated predictions are never wrong. Dated ones are, which is the point. Put the date in a calendar, not in the document.

    4. 04

      Name the one observation that would move you most

      This is the tripwire. It converts a static forecast into an instruction to watch something specific, and it is often more useful than the forecast itself.

    5. 05

      Score in batches, look for bias not verdicts

      Read twenty resolved cards at once. The finding is rarely 'we were wrong'; it is 'we are systematically overconfident above eighty per cent' or 'our timelines slip by one quarter, consistently.' Those are correctable.

    The shape of a filed card

    • Claim: at least one unaffiliated person uses the phrase 'vibe scientist' in a public bio or job description before 24 October 2026.
    • Probability: recorded before the window opened, not after.
    • Resolves: 24 October 2026, by search of public postings — method fixed in advance so the resolution cannot be argued.
    • Most-informative observation: a single unprompted use inside a company that is hiring, which would move the number further than ten social-media mentions.
    • Note: this card is live. Its resolution will be published whether or not it flatters the forecast.

    Built on

    Where it breaks

    Calibration requires resolution, so the method is blind to decisions whose outcomes never become observable — most organisational and hiring calls among them. It also degrades under observation: cards written to be scored by a manager drift toward safe claims near fifty per cent, which teaches nobody anything.

    Cite thisVibe Scientist. “The Calibration Card”, v1.0. https://www.vibescientist.com/frameworks#calibration-card. Retrieved July 29, 2026.

    The DRIFT Check

    v0.9 · draft

    Data-age · Retrieval · Inference · Framing · Trace

    A DRIFT check is a five-question audit run on any machine-produced synthesis before it is allowed to count as evidence: how old is the Data behind it, what Retrieval path selected the inputs, which Inference steps were the model's rather than the source's, whose Framing shaped the prompt, and can any single claim be Traced to a primary source you could open. A synthesis that fails the trace question is not evidence; it is a confident summary of an unknown population.

    Use when

    Any time an AI system clusters, summarises, or ranks the raw material of a decision — interview transcripts, tickets, market scans, literature. The check takes minutes; the belief it prevents can last years.

    1. D

      Data-age — what period does this actually describe?

      Models answer in the present tense about a corpus with a cut-off, and retrieval blends fresh and stale sources without marking the seam. Ask what window the underlying material covers.

    2. R

      Retrieval — what got in, and what could not?

      Every synthesis is a sample. Transcripts from one channel, documents one team happened to upload, pages a search index ranks well: the shape of the input is the shape of the conclusion.

    3. I

      Inference — which claims did the model add?

      Separate what a source says from what the system concluded across sources. Cross-source inference is where the useful ideas and the fabrications both live.

    4. F

      Framing — whose question was this?

      A prompt asking why customers love a feature will return reasons they love it. Re-run the synthesis with the opposite framing; the delta between the two answers is a better artefact than either.

    5. T

      Trace — open one primary source at random

      Pick a claim, follow it to the verbatim sentence, read it in context. One spot-check per synthesis is enough to keep the practice honest, and it fails more often than people expect.

    Running the check on a routine synthesis

    • A model clusters ninety support conversations into six themes with confident headline percentages.
    • Data-age: the export covers eleven weeks, seven of which preceded a pricing change — so two themes describe a product that no longer exists.
    • Retrieval: support conversations only. This is a sample of people who write to support, not of users.
    • Inference: the ranking of themes by importance was the model's; no source ranked anything.
    • Framing: the prompt asked for 'top user frustrations,' which guarantees six frustrations even in a happy population.
    • Trace: one theme's supporting quote turns out to be a paraphrase of two separate messages. The theme survives; the percentage does not.

    Built on

    Where it breaks

    Published as draft: the five questions come from repeated practice, not from a controlled comparison against un-audited synthesis. The check also cannot detect confident omission — what the corpus never contained will not announce itself, and no amount of tracing recovers a missing population.

    Cite thisVibe Scientist. “The DRIFT Check (Data-age · Retrieval · Inference · Framing · Trace)”, v0.9. https://www.vibescientist.com/frameworks#drift-check. Retrieved July 29, 2026.