THE COMPANY-BUILDING FIELD NOTEBOOKRESEARCH EDITION / SEPTEMBER 2026
Startup
Research.
Search
RESEARCH ADDITIONS / RESEARCH NOTE

An AI demo is not a deployment plan

A practical way to separate capability, failure handling and the cost of operating an AI workflow.

THE USEFUL DISTINCTION

Evaluate the whole workflow—including the person correcting it—not just the most impressive output.

Original paper-sculpture illustration of connected workspaces surrounding a layered software system
Original AI-generated editorial illustration · Not a documentary image

What the source establishes

NIST’s July 2024 Generative AI Profile is a voluntary, cross-sector companion to its AI Risk Management Framework. It organizes risks and potential actions; it does not certify that a particular product is safe. Its examples include confidently incorrect output, privacy problems, unsuitable human reliance and risks inherited from interconnected components. The document highlights governance, content provenance, pre-deployment testing and incident disclosure as important considerations. These are broader than a benchmark score. Read the NIST publication record and the profile itself.

This note uses that framework as a starting point. The workflow questions below are our editorial synthesis, not a NIST-approved test suite or evidence that any named vendor passes it.

Start with a bounded job

Write the task as an observable change in someone’s work. “Summarize these support conversations into a draft reviewed by an agent” is more testable than “automate customer service.” Name the inputs, the person who accepts the result, the system receiving it and the action that must never happen without approval.

Then collect examples that represent the intended job. Include ordinary requests, incomplete inputs, conflicting information and cases the system should decline. Keep a separate set for checking changes, so repeatedly improving against the same examples does not become the only definition of progress. Document whose work is absent from the sample.

For every example, decide what an acceptable result looks like before running the product. A plausible paragraph, a correct answer and an answer suitable for this customer’s situation are different judgments. Record disagreements between reviewers instead of hiding them inside one average.

Price the correction loop

Our proposed operating worksheet has five columns: task volume, accepted outputs, reviewer time, correction time and direct processing cost. Keep these quantities separate before attempting a per-task estimate. A fast first draft can still be a poor workflow if checking it requires recreating the work.

Ask whether the reviewer has enough context to spot an error. Who can stop the process? What happens when the underlying model, retrieval source or integration changes? Which failures require immediate escalation, and which can be corrected in the next review cycle? These questions turn “human oversight” from a reassuring phrase into an assigned responsibility.

Define the decision, not just the demo

A pilot should end with a documented choice: expand, narrow, redesign or stop. Choose the acceptance conditions, evaluation period and rollback owner beforehand. Treat a small successful pilot as evidence about that pilot, not every department or every language.

Our suggested evidence packet is deliberately modest: the task definition, sample limitations, evaluation results, correction costs, unresolved incidents and the next test. The goal is a decision someone else can inspect. This is a product-research worksheet, not a claim that a checklist eliminates model risk.

Inspect the source record.

Primary documents can establish what an organization reported or a regulator published. They do not independently prove every company claim. Our interpretation is labeled in the text.

  1. Generative Artificial Intelligence Profile — publication record

    NIST · Source date: 26 Jul 2024
    Retrieved: 16 Sept 2026

    What this source supports
    • Publication date; voluntary companion profile.
  2. NIST AI 600-1: Generative Artificial Intelligence Profile

    NIST · Source date: 26 Jul 2024
    Retrieved: 16 Sept 2026

    What this source supports
    • Risk examples and four primary considerations in the introduction; not product certification.

Edition & limitations

This is a bounded research note, not a comprehensive review or professional advice. Prepared 16 Sept 2026, revision 1. For this local review edition, no website publication date has been assigned. The dates above identify events and sources.

Read the evidence and image policy →
← All additions