THE COMPANY-BUILDING FIELD NOTEBOOKRESEARCH EDITION / SEPTEMBER 2026
Startup
Research.
Search
FIND EVIDENCE / PRACTICAL GUIDE

Run a concierge test without hiding the labor

Manual delivery can expose the real workflow, but only if the customer knows what is manual and the founder counts every minute required to deliver it.

  1. Manual delivery
  2. Measured effort
  3. Repeatable process
Conceptual relationship map, not measured data or a guaranteed sequence.

Concierge delivery is a good way to discover what the product must do. It is also an excellent way to fool yourself. A founder performs judgment at zero recorded cost, fixes exceptions in private, and reports that customers received the outcome. Demand may be real while the proposed business is not.

The source chapter distinguishes a disclosed concierge service from hidden manual work and warns that founder labor can make bad economics look healthy. Read it at validation chapter.

State what the customer is buying

Tell the customer which steps are performed by people, which are assisted by software, expected turnaround, review rights, data handling, and what may change during the test. Do not present a human-produced result as autonomous output.

FTC small-business guidance says advertising claims must be truthful and substantiated, and that a material omission can mislead; a necessary qualification should be clear and close to the claim (FTC advertising FAQ). That is US regulator guidance, not legal advice for every jurisdiction. Editorially, disclosure is also good research design: when users know where judgment sits, their feedback distinguishes trust in the outcome from belief in an automation claim.

Instrument the service like a product

For every job, log:

  • intake and clarification minutes;
  • execution minutes by task;
  • quality review and rework;
  • exceptions and judgment calls;
  • tools or third-party costs;
  • customer support;
  • elapsed turnaround;
  • outcome accepted, revised, or rejected.

Paul Graham's essay recommends manually recruiting and serving early users, including using software on their behalf, because the work creates a fast learning loop (Do Things that Don't Scale). It is a founder-investor essay and an argument from experience, not a representative study. Keep the learning loop; do not import the implication that labor eventually disappears.

Worked hypothetical example

Hypothetical: A founder offers weekly inventory recommendations to five independent shops. Customers know that the dashboard is a prototype and that a person reviews imports and writes the recommendation. Each shop pays $300 for a four-week test.

The four-week fee allocates $75 per week. In week one, a shop requires 35 minutes of data cleanup, 25 minutes of analysis, 15 minutes of review, and 10 minutes of support: 85 minutes. At an internal labor cost assumption of $60 an hour, labor is $85—already $10 above the weekly allocation before software and payment costs. If that week repeats four times, labor reaches $340 against $300 of revenue. The ledger reveals that data cleanup is the largest repeated step and unusual supplier codes cause most rework.

The next test automates only normalization, with a stop rule: if cleanup does not fall below 12 minutes for four of five shops, the founder will not build the forecasting layer. All figures are fictional and demonstrate method, not a benchmark.

Classify each manual step

At the end of each week, put work into four buckets:

  • Automate: repeated, rule-like, and worth engineering.
  • Standardize: human work that becomes faster with templates or constraints.
  • Charge for: expert judgment customers value and software will not remove soon.
  • Stop: exceptions that do not support the target segment.

This is editorial analysis. The important move is to avoid labeling every human step 'future automation.' Some steps are the service; others are warnings that the segment or promise is too broad.

Exit checklist

  • Customers received the promised disclosure and can ask for human review.
  • Labor is timed by activity, including rework and support.
  • Quality failures and edge cases are retained, not smoothed out.
  • The next automation target has a measurable labor or quality objective.
  • The founder has tested whether customers will pay when labor is priced honestly.
  • There is a cap on customers or jobs until the next decision.

Limits

Manual delivery can establish that an outcome matters and reveal a workflow. It does not show that software can reproduce expert judgment, that quality holds at volume, or that support load will remain stable. Privacy, employment, professional-services, and sector rules may apply. Get appropriate advice where the manual work touches regulated decisions or sensitive data.

Sources & scope

Sources checked 19 September 2026. Worked scenarios are illustrative; recommendations are editorial analysis. These checks do not re-verify the entire original notebook.

  1. Advertising FAQ's: A Guide for Small Business — U.S. Federal Trade Commission

    FTC guidance says material express and implied claims require support and qualifying disclosures must be clear and conspicuous.

    Source publication date: Not established · Retrieved 2026-09-19

  2. Do Things that Don't Scale — Paul Graham

    Graham argues that founders often need to recruit users manually and sometimes use the product on customers' behalf to learn the workflow.

    Source publication date: Not established · Retrieved 2026-09-19

Developed from the original notebook

Keep the question moving.

Next in this path: Design a paid pilot that can end cleanly

All practical guides →