THE COMPANY-BUILDING FIELD NOTEBOOKRESEARCH EDITION / SEPTEMBER 2026
Startup
Research.
Search
RESEARCH LIBRARY / Company economics

Know what the numbers mean

Metric definitions, worked examples, benchmark limitations, and eight illustrative financial models.

Interactive / 02 · Cash runway

See the cash story
month by month.

Change the assumptions to trace a simple 24-month cash path. It is a bounded illustration, not a forecast or financial advice. Read the runway concept in the research chapter.

Fictional cash scenario assumptions

Fictional defaults. Calculated in your browser; inputs are not saved or sent. Month 1 uses the receipt amount entered above; growth changes receipts from month 2 onward. Spend stays fixed.

Underlying monthly values (USD)
Illustrative monthly cash balance values in United States dollars
MonthReceipts (USD)Spend (USD)Net (USD)Cash available (floored at zero, USD)Unfunded shortfall (USD)

Amounts are rounded to cents each modeled month. This simple arithmetic does not model taxes, financing, collection timing, staffing changes, or borrowing. A zero balance counts as exhausted; an unfunded shortfall is shown without assuming debt.

Research date: September 15, 2026. Every benchmark below is dated. Figures move; the reasoning does not.

How to read this chapter

Part A defines the metrics and shows how to compute them. Part B builds eight illustrative financial models. Parts C and D cover which metrics matter when, and the ways founders routinely fool themselves with numbers.

Six things to hold in mind throughout:

1. Medians, not means. Startup outcomes are power-law distributed, so averages are dragged upward by a handful of extreme performers. Where a source reports a mean I say so explicitly. Otherwise assume median.

2. Survivorship bias is severe, and it is not a footnote. Almost every benchmark panel in this chapter samples companies that (a) still exist, (b) have enough revenue to be worth benchmarking, (c) use a financial system that reports data, and (d) chose to respond to a survey. The Aleph × Benchmarkit panel of 342 companies, the SaaS Capital panel of 1,000+, the High Alpha panel of 800+ — these are cohorts of survivors, and often of funded survivors. Companies that died before hitting $1M ARR are absent. Companies that quietly stagnated are underrepresented because stagnating founders do not fill out benchmarking surveys. The practical consequence: the true median company in your category is worse than the median in any published benchmark. If your numbers sit at the 25th percentile of a published panel, you are probably around the median of all companies that ever tried.

3. Self-reported data is unaudited. Most benchmark surveys ask founders and CFOs to report their own numbers, with the reporting company's own metric definitions. "NRR" means five different things across five respondents. Treat ±3-5 points as noise.

4. The models in Part B are illustrative, not forecasts. They are arithmetic exercises built on stated, cited assumptions to show how the parts interact and where they break. They are not predictions about any real company, and they are not advice. Each one includes a "typical/median" path alongside the "if it works" path, because the second is what pitch decks show and the first is what usually happens.

5. Nothing here is personalized financial, investment, tax or accounting advice. Revenue recognition in particular has real accounting consequences; the treatment described here is the conceptual version, and an actual company needs an actual accountant.

6. Metrics are instruments, not goals. Every metric in this chapter has a documented failure mode where optimizing the number destroys the thing the number was measuring. Those failure modes are listed for each metric, because in practice they cause more damage than not measuring at all.


PART A — THE METRICS

A0. Three things to fix before any metric means anything

Before computing anything, settle three definitional questions, write the answers down, and do not change them without noting the change.

What is a "customer"? A billing account, a legal entity, a seat, or a workspace? A company with 40 departments buying separately is 40 customers by one definition and one by another. Your logo churn, ARPA, and CAC all change by multiples depending on the answer.

What counts as "revenue" for a given month? Cash received, amount invoiced, or amount earned? These are three different numbers (see A1).

Which costs are "acquisition" costs? Just paid media, or the whole go-to-market org? A 3x difference in reported CAC is common between the narrowest and broadest definitions of the same company.

Almost every metric argument between a founder and an investor is actually an argument about one of these three definitions.


A1. Revenue: bookings, billings, recognized revenue, MRR, ARR

These four numbers describe the same commercial activity at four different moments, and confusing them is the single most common source of startup accounting error.

A1.1 Bookings

Definition. The total contract value a customer has committed to, at the moment they sign. It is a sales metric, not an accounting metric — it appears in no financial statement.

Formula. Bookings = Σ (total committed contract value of deals signed in period)

Worked example. In March you sign three deals: a 2-year contract at $30,000/year paid annually; a 1-year contract at $24,000 paid monthly; and a 3-year contract at $50,000/year with a 6-month free period at the start. March bookings = $60,000 + $24,000 + $150,000 = $234,000. Note that $150,000 of that is a commitment stretching to 2029, and $25,000 of it will never be billed because of the free period.

What good looks like. Bookings should grow faster than revenue in a healthy expanding company (you are signing longer and larger contracts than you are recognizing). If bookings and revenue grow at the same rate, contract terms are static. If bookings grow much faster than billings, you are signing deals with long ramps or deferred starts.

How it gets gamed or misread. Bookings is the most inflatable revenue number in a startup, because nothing constrains it. Multi-year contracts inflate the headline; a three-year deal is booked once at 3x annual value. "Total contract value" that includes optional renewal years, unexercised expansion options, or non-binding letters of intent is not a booking. The classic tell: a company reporting "$10M in bookings" that collects $2M in cash. Ask for annual contract value (ACV), not total contract value (TCV), and ask what the weighted average contract length is.

A1.2 Billings

Definition. What you actually invoiced in the period. This is the number most closely tied to cash.

Formula. A common public-company approximation: Billings = Recognized revenue + Change in deferred revenue

Worked example. In Q2 you recognize $500,000 of revenue. Deferred revenue rises from $1.2M to $1.5M. Billings = $500,000 + $300,000 = $800,000. You invoiced $800,000 but only earned $500,000; the $300,000 difference is money you hold but have not yet delivered against.

What good looks like. For a subscription business selling annual prepaid contracts, billings run meaningfully ahead of recognized revenue during growth. Deferred revenue growing as fast as or faster than ARR is a healthy sign in an annual-prepay business.

How it gets gamed or misread. Billings can be pulled forward by offering discounts for annual or multi-year prepay at quarter-end, which flatters cash and billings while reducing lifetime revenue. Conversely, a company shifting from annual prepay to monthly billing will show collapsing billings and deferred revenue while its actual business is fine. Always ask whether the billing mix changed.

A1.3 Recognized revenue

Definition. Revenue earned in the period, under accrual accounting (in the US, ASC 606). For a subscription, you recognize ratably as you deliver the service.

Formula. For a straightforward subscription: Monthly recognized revenue = Annual contract value ÷ 12, recognized as service is delivered.

Worked example. A customer signs a $120,000 one-year contract on April 1 and pays the full amount up front. April bookings: $120,000. April billings: $120,000. April recognized revenue: $10,000. The other $110,000 sits on the balance sheet as deferred revenue (a liability — you owe eleven months of service) and unwinds at $10,000/month.

The gross-vs-net trap. ASC 606 also determines whether you book the gross amount a customer pays or only your net cut. If you control the good or service before transferring it, you are a "principal" and book gross. If you merely arrange for another party to provide it, you are an "agent" and book only your commission. This distinction is central to marketplaces — and, newly, to AI companies that resell model capacity. CJ Gustafson's analysis of the AI ARR question walks through the case of Mercor, which was widely reported at ~$500M ARR; at a roughly 30% take rate, the net revenue figure would be closer to ~$150M, with the rest passing through to the contractors doing the work (Mostly Metrics, September 28, 2025). Whether a given company is principal or agent is a genuine accounting judgment, not a slogan — but a founder should know which one they are.

How it gets gamed or misread. Recognizing setup fees, implementation services and hardware sales up front while calling the whole thing "recurring." Recognizing a full year on signature. Presenting gross flow-through as revenue. And, most common in early-stage startups, simply not doing accrual accounting at all and calling cash collections "revenue" — which makes a company look lumpy and unpredictable when it is neither.

A1.4 MRR and ARR

Definition. Monthly Recurring Revenue is the normalized, repeatable subscription revenue in a month. ARR is the annualized version. Both deliberately exclude one-time revenue.

Formula. MRR = Σ (normalized monthly subscription value of active customers); ARR = MRR × 12

MRR is usefully decomposed: Ending MRR = Starting MRR + New MRR + Expansion MRR − Contraction MRR − Churned MRR

Worked example. You start January with $50,000 MRR. You add $6,000 from new customers, $2,000 from upgrades, lose $1,500 to downgrades, and $3,000 to cancellations. Ending MRR = 50,000 + 6,000 + 2,000 − 1,500 − 3,000 = $53,500. Net new MRR = $3,500. ARR = $642,000. Note that gross new MRR of $8,000 became net new MRR of $3,500 — 56% of your new sales went to replacing losses.

What good looks like. Growth rate is the headline. Across 1,000+ private B2B SaaS companies surveyed in March 2026, median growth was 22% in 2025, down from 25% in 2024; bootstrapped companies grew at a 20% median and equity-backed at 25% (SaaS Capital, 2026 growth benchmarks). 7.3% of respondents were flat or shrinking. Note what this means: the median venture-backed private SaaS company in 2025 grew 25%, not 100%. Growth rates in the triple digits are a small top slice, not a norm.

By ARR band, the 2025 High Alpha panel (800+ companies, surveyed Aug–Sep 2025) reported median growth of roughly 40% at $1–5M ARR, 30% at $5–20M, and 15% above $50M (Growth Unhinged summary of the 2025 SaaS Benchmarks Report).

How it gets gamed or misread. Annualizing the best month. Including one-time services revenue, hardware, or professional services in "ARR." Including customers in a free trial at their list price. Including signed-but-not-started contracts. Including revenue from customers who have given notice. Each of these is individually defensible-sounding and collectively produces an ARR figure 20-40% above the real recurring base.

A1.5 Why ARR is misleading for usage-based and AI businesses

ARR was invented for a specific world: a customer signs a contract for a fixed number of seats at a fixed price for twelve months, and unless they cancel, that revenue repeats. In that world, annualizing a month is a reasonable prediction.

Usage-based and AI-native businesses break three of those assumptions at once.

The revenue is not contracted, it is consumed. If a customer's spend is a function of how much they use, this month's revenue is an observation, not a commitment. Annualizing it assumes usage holds — but usage is volatile, seasonal, and often driven by a customer's own experimental budget rather than a production dependency.

Early AI revenue is disproportionately experimental. A large share of 2024–2026 AI application revenue came from enterprise innovation budgets — pilots funded to answer "should we be doing this?" rather than to run a business process. Pilot revenue annualizes beautifully and renews poorly.

The "R" may not be recurring at all. Where the company is effectively reselling model capacity and taking a cut, the annualized gross flow is not the company's revenue in any meaningful sense (see the principal/agent point in A1.3).

There is an active debate about replacement metrics — "Annual Predictable Revenue," "Annual Consumption Run-Rate," and similar — and none has won. What is clear is what a careful reader should demand instead of a single ARR number:

  • Committed vs. uncommitted split. How much of the run-rate is contractual minimum commitment, and how much is discretionary usage above it?
  • Trailing twelve months of actual revenue, not an annualized month.
  • Net revenue retention on a fixed cohort, which captures whether usage actually expands (see A7).
  • Gross vs. net revenue, stated explicitly.
  • Concentration. What share of run-rate comes from the top five accounts?

Benchmark panels are starting to separate this out. In the 2026 Aleph × Benchmarkit panel (342 B2B SaaS and AI-native companies, full-year 2025 actuals, published June 2026), usage-based companies showed a median NRR of 108% versus 98% for seat-based — but with a far wider spread, with the 75th percentile of usage-based companies at 155% (Aleph, NRR benchmarks 2026). High variance in both directions is the signature of usage-based revenue: it compounds hard when it works and evaporates when it doesn't.


A2. Gross margin, and what belongs in COGS for software

Definition. Gross margin is the share of revenue left after the direct cost of delivering the product. It is the ceiling on every other margin in the business and the single most important structural number in a company's P&L.

Formula. Gross margin % = (Revenue − COGS) ÷ Revenue × 100

What belongs in software COGS. The test is variability with delivery: if the cost rises when you serve one more customer or one more unit of usage, it is COGS. The conventional software COGS line items:

  • Hosting and cloud infrastructure for the production product
  • Third-party APIs and data the product calls to function
  • Customer support (the tier that keeps existing customers working)
  • Customer success salaries allocated to service delivery, not to expansion selling
  • Payment processing fees
  • Software licenses embedded in the product
  • Amortization of capitalized software, where applicable

Not COGS: R&D, sales, marketing, general and administrative, and the cloud spend for development and staging environments.

Worked example. A company with $2,000,000 in revenue has $180,000 of production hosting, $60,000 of third-party APIs, $150,000 of support salaries, $50,000 of payment fees. COGS = $440,000. Gross margin = (2,000,000 − 440,000) ÷ 2,000,000 = 78%.

The 2025–2026 question: do inference costs belong in COGS?

Yes — if the inference happens to serve a customer request. This is not really controversial in accounting terms; inference is a direct, variable cost of delivering the product, exactly like a third-party API call. The reason it feels controversial is that recognizing it honestly makes the margins look worse, and a generation of founders and investors had internalized 80% gross margin as a law of nature.

Three practical splits matter:

  1. Inference serving customer requests → COGS. Every token generated in response to a user action.
  2. Training, fine-tuning and evaluation → R&D. These are investments in the product, not costs of delivering it, and are properly operating expenses (capitalization questions aside).
  3. Internal use of AI tools → operating expense in the relevant department. Your engineers' coding assistant subscriptions are R&D, not COGS.

The trap that catches companies is item 2 sliding into item 1. If your product requires per-customer fine-tuning to work, that fine-tuning is arguably a delivery cost, and the margin profile of the business is much worse than the headline suggests.

A useful derived metric. Ben Murray's inference efficiency ratioAI product revenue ÷ inference cost — makes the exposure explicit. His suggested targets: roughly 10:1 for AI-infused SaaS (inference ≈ 10% of AI product revenue) and roughly 5:1 for AI-native products (inference ≈ 20%) (The SaaS CFO, May 2026). Below 3:1, the business is not a software business in the margin sense.

The benchmarks

Traditional SaaS, 2025 actuals. Median software gross margin 80%; median total-revenue gross margin 76% (total revenue includes services); top quartile 86%+; bottom quartile 50%. Companies on usage-only pricing showed 62% (Aleph × Benchmarkit 2026, n=228–232 of 342). The Benchmarkit 2025 report (CY2024 data) showed a similar picture: 77% blended, 81% subscription-only, 30% on professional services (Benchmarkit 2025 SaaS Performance Metrics).

The compression question is genuinely unsettled. The Aleph × Benchmarkit authors explicitly report that software gross margin "held remarkably stable at 79–81% across all four years" of their series and that they could not prove an AI-margin-compression hypothesis at the median. The High Alpha 2025 panel, by contrast, found gross margins for AI-native startups running about 5 points below traditional SaaS peers, and reported roughly 10 points of year-over-year compression among early-stage companies (Growth Unhinged, 2025 SaaS Benchmarks).

Both can be true. The median of a panel dominated by established subscription businesses barely moves, while the subset of young, inference-heavy products sits far below it. Look at the subset you actually belong to, not the panel median.

AI-native margins specifically. Bessemer's segmentation of fast-growing AI companies is the most striking data point. Their "AI Supernovas" — companies reaching very large revenue in year one — carried average gross margins around 25%, and often negative, while their "Shooting Stars" cohort averaged 60% (Bessemer, The State of AI 2025, August 13, 2025). These are means across small cohorts of extreme outliers and should not be read as typical, but they establish the range: the fastest-growing AI revenue in the market has been, in part, subsidized revenue.

The counterargument, stated fairly. a16z's Sarah Wang and Martin Casado argue that low margins at a point in time say little about the long-run model: companies move from unlimited introductory pricing to tiered and usage-based pricing, rate-limit the unprofitable top few percent of users, benefit from competition among model providers, and ride a cost curve that has fallen "anywhere from 10x to 100x+ in the last 18 months" (a16z, August 21, 2025).

The counter-counterargument. Cost per token falls; cost per task often does not, because customers migrate to newer, more expensive models and to reasoning modes that consume far more tokens per request. As Gustafson put it, older cheap models are "as desirable as a flip phone at an iPhone launch." As of September 2026, frontier-tier models are priced around $5–10 per million input tokens and $25–50 per million output tokens, while production-grade smaller models sit near $0.75/$3.75 and budget models run as low as $0.03–$0.28 input (BenchLM LLM pricing, September 2026). The spread between tiers is two to three orders of magnitude — which means your margin is mostly a routing and product design decision, not a market condition.

How gross margin gets gamed or misread. Classifying support as G&A. Classifying production hosting as R&D. Excluding free-tier infrastructure costs on the theory that free users are marketing. Reporting "subscription gross margin" while services are 30% of revenue at 30% margin. Netting cloud credits (which expire) against COGS without disclosure — a very common early-stage distortion, since a company burning $40,000/month of a $200,000 cloud credit shows a gross margin that will collapse in month six.


A3. Burn rate and runway

Definitions.

  • Gross burn = total cash operating outflows in a month, ignoring revenue.
  • Net burn = gross burn minus cash collections. This is the rate at which the bank balance actually falls.
  • Runway = months until cash reaches zero at the current net burn.

Formulas. Gross burn = total cash operating expenses in month Net burn = gross burn − cash received in month Runway (months) = cash on hand ÷ average net burn over the trailing 3 months

Worked example. A company holds $1,800,000. Monthly cash operating costs are $260,000. Monthly cash collections are $95,000. Gross burn = $260,000. Net burn = $165,000. Runway = 1,800,000 ÷ 165,000 = 10.9 months.

But the company is growing collections 8% per month. Modelling that forward, collections reach ~$190,000 in month 9 while costs hold flat, so net burn falls over time and real runway is closer to 13–14 months. This is why a single-point runway number is nearly always wrong: runway must be modelled, not divided. Model it three ways — flat plan, plan minus 30% on revenue, and plan with hiring frozen — and quote the middle one.

What good looks like. The durable convention is that you want 18 months of runway after a round closes, and you should start raising with 9–12 months left, because a raise takes three to six months and you cannot negotiate from a position of two months of cash. In a market where down rounds are historically rare — Carta reported a down-round rate of 11.4% in Q1 2026, a return to pre-pandemic levels (Carta, State of Private Markets Q1 2026, published May 29, 2026) — this convention has loosened somewhat, but the arithmetic of a raise taking a quarter or two has not changed.

How burn gets gamed or misread. Quoting gross burn as net burn (or vice versa) depending on which is flattering. Excluding one-time costs from "normalized burn" every single month — if you have an unusual expense every month, it is not unusual. Timing payables to push spend into the next month before a board meeting. Counting an unsigned term sheet or an undrawn debt facility in "cash on hand." Ignoring that annual prepayments make collections lumpy, so a single month's net burn can be wildly unrepresentative — always use a trailing three-month average. And the largest one: calculating runway on a flat cost base while the plan assumes hiring. If the plan adds eight people, runway on today's burn is fiction.


A4. Customer acquisition cost (CAC)

Definition. The cost to acquire one new paying customer. There are three common versions and they can differ by 3x for the same company.

Formulas.

  • Blended CAC = total sales & marketing spend ÷ all new customers (including organic/referral)
  • Paid CAC = paid acquisition spend ÷ customers attributable to paid channels
  • Fully-loaded CAC = (all S&M salaries + commissions + tools + programs + paid media + allocated overhead) ÷ new customers

Worked example. In a quarter, a company spends $90,000 on paid media, $150,000 on sales and marketing salaries, $20,000 on tools and events. It acquires 200 new customers, of which 70 are attributable to paid channels.

  • Blended CAC = ($90,000 + $150,000 + $20,000) ÷ 200 = $1,300
  • Paid CAC = $90,000 ÷ 70 = $1,286
  • Fully-loaded CAC — same as blended here, since we included everything — $1,300
  • A narrow "marketing CAC" that counts only paid media against all customers = $90,000 ÷ 200 = $450

That last number is the one that appears in optimistic decks. It is not wrong so much as meaningless: it credits paid spend with customers it did not acquire.

Which to use when. Use fully-loaded blended CAC for company-level economics and board reporting — it is the only version that cannot be gamed by attribution choices. Use paid CAC by channel for allocating marketing budget, because it answers the marginal question: if I spend another dollar here, what happens? Report both. Never report only the second.

What good looks like. CAC in absolute dollars is not comparable across companies; the ratio versions are. The CAC ratio — dollars of S&M spent per dollar of new ARR acquired — had a median of $2.00 for new-customer acquisition and $1.00 for expansion in the Benchmarkit 2025 report (CY2024 data), with the new-customer figure up 14% year over year (Benchmarkit 2025). Expansion revenue costing half as much to acquire as new revenue is the standard structural finding, and is the whole argument for investing in retention and account growth.

For spending levels, SaaS Capital's March 2026 survey of 1,000+ private B2B SaaS companies found medians of 15% of ARR on sales, 8% on marketing, 9% on customer support/success, 22% on R&D and 15% on G&A. Equity-backed companies spent roughly 70% more on sales and 100% more on marketing than bootstrapped peers; total median spend was 96% of ARR for bootstrapped and 101% for equity-backed (SaaS Capital, 2026 Spending Benchmarks).

How CAC gets gamed or misread. Excluding salaries — this is the most common single distortion and typically halves reported CAC. Excluding founder time in the early stage, which makes the first 50 customers look nearly free and the next 200 look like a catastrophe. Crediting organic and word-of-mouth customers against paid spend. Counting free-trial signups as "customers." Computing CAC in the same period as the spend, when the spend that produced this quarter's customers happened last quarter. Averaging across wildly different segments, so a $400 self-serve CAC and a $60,000 enterprise CAC blend into a meaningless $3,000.

The deeper issue: CAC is not a constant. It is a supply curve. The first customers are cheap because they are the most motivated; each additional increment of volume costs more. A company that reports "our CAC is $1,200" and models 10x growth at $1,200 has assumed away the central difficulty of scaling. Model CAC as rising with volume — a common working assumption is that CAC rises 15–30% for each doubling of paid volume within a channel, though the real number is empirical and channel-specific.


A5. Lifetime value (LTV), and why the classic formula breaks

Definition. The total gross profit a customer generates over their relationship with you, discounted to present value.

The classic formula. LTV = (ARPA × Gross margin %) ÷ Churn rate

Worked example. ARPA is $200/month, gross margin 80%, monthly churn 3%. LTV = (200 × 0.80) ÷ 0.03 = $5,333. With CAC of $1,200, LTV:CAC = 4.4x. This looks excellent.

Now change monthly churn from 3% to 2%. LTV = 160 ÷ 0.02 = $8,000, and LTV:CAC = 6.7x. A one-point change in an input you measured with noise moved the headline output by 50%. Change it to 4% and LTV:CAC falls to 3.3x. The formula is a division by a small, uncertain number, which means it is arithmetically unstable in exactly the region where startups operate.

Why the formula breaks

Bill Gurley's 2012 essay remains the definitive critique, and every point in it has aged well (Above the Crowd, September 4, 2012). The failures worth internalizing:

1. Dividing by churn assumes churn is constant forever. It never is. Real churn is high early and declines as a cohort seasons — the customers who were going to leave leave first. Using a blended current churn rate to extrapolate a lifetime overstates the value of new cohorts and understates the value of old ones.

2. It assumes an infinite time horizon. 1 ÷ 0.02 implies a 50-month average life, and the tail extends for years. A startup that is three years old has no evidence about what happens in year seven. Discounting helps but does not fix the fact that you are extrapolating past the edge of your data.

3. The variables are interdependent — Gurley's most important point. Raising price raises ARPA and raises churn. Increasing marketing spend raises CAC and often lowers the quality of acquired customers, which raises churn, which lowers LTV. The formula treats these as independent knobs. They are not. This is why "we'll just raise prices to fix LTV:CAC" so often fails.

4. Purchased customers underperform organic ones. Cohorts acquired through paid channels systematically retain worse and spend less than cohorts that arrived by word of mouth. A blended LTV computed across both, then compared to a paid CAC, is comparing the wrong things.

5. It ignores competitive dynamics. A high LTV is an invitation for competitors to spend more on acquisition than you do. LTV is a measurement, not a moat.

6. It is used to rationalize. Gurley's blunt framing: executives use LTV to justify marketing budgets they have already decided on. When the LTV model and the spend decision disagree, it is usually the model that gets revised.

What to use instead

Cohort-based payback and cumulative gross profit. For each monthly acquisition cohort, chart cumulative gross profit per customer against cumulative acquisition cost. Read off the month where the lines cross (that is your real payback period), and read off the observed cumulative value at 12, 24 and 36 months. This uses only data you actually have.

Capped LTV. Compute LTV over a defined horizon — 24 or 36 months — rather than to infinity. It is more conservative, more comparable across companies, and far harder to inflate.

Report the retention curve, not the ratio. A retention curve that flattens is worth more than any LTV number, because a flat tail is the actual claim being made.

What good looks like, with heavy caveats. The widely-quoted target is LTV:CAC above 3x; one 2026 aggregation cites a 3.3x median for private SaaS, higher for bootstrapped companies and lower for enterprise sales-led motions (SaaSRise 2026 benchmark compilation). Treat that number as folklore with a citation. It is computed inconsistently everywhere and it is the least reliable metric in this chapter. CAC payback period (A9) measures the same underlying property with far less room for fantasy, which is why sophisticated investors ask for it instead.


A6. Churn: logo vs revenue, gross vs net

Definitions. Four distinct metrics, frequently conflated.

  • Logo (customer) churn — the share of customers lost.
  • Gross revenue churn — the share of revenue lost to cancellations and downgrades.
  • Net revenue churn — revenue lost minus revenue gained from existing customers via expansion.
  • Gross vs net is the expansion question; logo vs revenue is the "which customers" question.

Formulas. Logo churn = customers lost in period ÷ customers at start of period Gross revenue churn = (churned MRR + contraction MRR) ÷ starting MRR Net revenue churn = (churned MRR + contraction MRR − expansion MRR) ÷ starting MRR

Worked example. Start the month with 500 customers and $250,000 MRR. You lose 20 customers representing $6,000 MRR, suffer $4,000 of downgrades, and gain $9,000 of expansion from existing accounts.

  • Logo churn = 20 ÷ 500 = 4.0%
  • Gross revenue churn = (6,000 + 4,000) ÷ 250,000 = 4.0%
  • Net revenue churn = (10,000 − 9,000) ÷ 250,000 = 0.4%

Here logo churn equals gross revenue churn, meaning you are losing average-sized customers. If gross revenue churn were 2%, you would be losing your smallest customers — usually fine. If it were 7%, you would be losing your largest — an emergency, even though the logo number looks identical.

The annualization trap. Monthly churn does not annualize by multiplying by 12. Use Annual churn = 1 − (1 − monthly churn)^12. At 4% monthly, that is 1 − 0.96^12 = 38.7% annual, not 48%. At 2% monthly, 21.5% annual. Getting this wrong in either direction is common.

What good looks like. Annual gross revenue retention (the complement of gross revenue churn) had a median of 84% in 2025, down 4 points from 88% in 2024, with the top quartile at 91% and bottom quartile at 76%. The decline hit every segment. By deal size: 91% at $50K–$100K ACV, 80% at sub-$5K ACV. By motion: 88% sales-led, 80% hybrid, 79% product-led (Aleph × Benchmarkit 2026, n=226). The report attributes the drop to longer sales cycles, harder ROI scrutiny, and customers weighing incumbents against AI-native alternatives.

Bootstrapped companies in the $3–20M ARR range showed a median GRR of 91% and a 90th percentile of 100% (SaaS Capital, 2026 bootstrapped benchmarks, published April 24, 2026) — better than the broader panel, consistent with bootstrapped companies selling to customers who chose them rather than being sold to.

For consumer subscriptions the numbers are structurally different and much worse. Across 115,000+ apps and $16B of tracked revenue, RevenueCat found roughly 27–28% of subscribers still subscribed at twelve months, and that about 72% of annual subscriptions cancel within the first year, with 35% of all annual-plan cancellations occurring in month one (RevenueCat, State of Subscription Apps 2026, published March 19, 2026). A consumer subscription app losing three-quarters of its subscribers in a year is normal, not broken.

How churn gets gamed or misread. Excluding customers who churned in their first 90 days ("they weren't really customers"). Excluding annual contracts from monthly churn calculations, which hides the fact that annual customers churn in a lump at renewal. Measuring churn against a rapidly growing denominator, which mechanically suppresses the rate — if you doubled your customer count this month, this month's churn rate is meaningless. Reporting net churn as "churn" without saying so. Counting a downgrade from $10,000 to $500/year as retention. And the reverse error: obsessing over logo churn in a business where the top 10 accounts are 60% of revenue, where logo churn is nearly irrelevant.


A7. Net revenue retention (NRR / NDR)

Definition. How much revenue a fixed cohort of customers generates twelve months later, including expansion, contraction and churn, excluding any new customers. It is the single best measure of whether a product deepens or decays inside its installed base.

Formula. NRR = (Starting ARR + Expansion − Contraction − Churn) ÷ Starting ARR × 100

Measured on a fixed cohort over 12 months, with no new logos in the numerator.

Worked example. Twelve months ago, a cohort of customers represented $4,000,000 ARR. Over the year: $900,000 of expansion, $250,000 of contraction, $500,000 of churn. NRR = (4,000,000 + 900,000 − 250,000 − 500,000) ÷ 4,000,000 = 103.8%

The same cohort's GRR = (4,000,000 − 250,000 − 500,000) ÷ 4,000,000 = 81.3%. NRR above 100% with GRR in the low 80s means a small number of accounts are expanding hard while a large number leak away — a real but fragile position, and one that breaks the moment the expanding accounts mature.

Why NRR above 100% matters so much. At NRR of 110%, revenue grows 10% a year with zero new sales. At 90%, you must sell 10% of your base in new business just to stand still. This compounds, which is why NRR drives valuation more than almost any other operating metric.

What good looks like. Median NRR of 102% in 2025, with the 75th percentile at 110% and the 25th at 92%; best-in-class is 120%+. By pricing model, usage-based 108% vs seat-based 98%. By ARR band, 94% below $5M, 101% at $20–50M, 103% above $100M. By ACV, 105% in the $25K–$50K band, below 100% in the $10K–$25K band, 98% sub-$5K (Aleph × Benchmarkit 2026, n=230).

SaaS Capital's independent panel found median NRR of 103% for bootstrapped companies at $3–20M ARR (90th percentile 117.9%), and reported that moving NRR from the 90–100% band into the 100–110% band correlated with about 5 percentage points of additional growth (SaaS Capital 2026).

A crucial structural finding from ChartMogul's panel of 2,500+ SaaS businesses: NRR falls as customer count rises. Among companies with more than 12,000 subscribers, only about 6% achieved NRR at or above 100%. Companies with NRR ≥100% grew at a 48% median year over year; expansion made up more than half their revenue growth (ChartMogul, The SaaS Retention Report: The New Normal, H1 2024 data). The implication is that 120% NRR is essentially an enterprise, high-ACV phenomenon, and a self-serve SMB product should not plan around it.

How NRR gets gamed or misread. Including new logos in the numerator — the single most common inflation, and it can add 20 points. Choosing the cohort start date to exclude a bad quarter. Measuring only "active" customers, which drops the churned ones out of the denominator entirely and can push NRR above 130% on a shrinking business. Counting a one-time services upsell as recurring expansion. Reporting NRR without GRR, which hides whether you are retaining a base or riding a few whales. Reporting quarterly NRR annualized, which compounds noise. Always ask for GRR alongside NRR. The gap between them is the most informative number in the pair.


A8. Conversion rates

Definition. The share of people who move from one step of a funnel to the next. The metric is trivial; the errors are all in denominators and time windows.

Formula. Conversion rate = entities completing step N+1 ÷ entities entering step N, within a defined time window, measured on a cohort.

Worked example. In March, 10,000 visitors reach your pricing page; 620 start a free trial; 112 convert to paid within 30 days of trial start; 96 are still paying at day 90.

  • Visitor → trial: 6.2%
  • Trial → paid (30-day window): 18.1%
  • Visitor → paid: 1.12%
  • 90-day retained conversion: 0.96%

The last number is the one that matters for economics, and it is 15% lower than the headline conversion rate.

What good looks like.

B2B self-serve. One 2026 compilation of ChartMogul data puts trial-to-paid at roughly 5.6% for freemium, 8.9% for opt-in trials without a credit card, and 31.4% for credit-card-required trials (SaaSRise 2026 compilation). The pattern — requiring a card multiplies conversion several-fold while shrinking trial volume — is consistent across sources.

Sales-led B2B. The same compilation cites demo-to-close rates of about 32% SMB, 25% mid-market, 18% enterprise.

Consumer subscription apps. The RevenueCat 2026 data is the best-sourced set available. Median day-35 trial-to-paid conversion was 10.7% for hard paywalls versus 2.1% for freemium. Longer trials (17–32 days) converted at 42.5% versus 25.5% for trials under four days — but note this is conversion of trial starters, and long trials attract fewer starters. 55.4% of trial cancellations happen on day zero. Realized revenue per install at day 60 was $3.09 for hard paywalls versus $0.38 for freemium (RevenueCat, March 2026).

Activation. Across 500+ responses to Lenny Rachitsky's benchmarking survey, median activation rate was 25% (mean 34%); for SaaS specifically, median 30% (Lenny's Newsletter, October 25, 2022). This is older data but activation definitions are so product-specific that newer panel medians would not be more useful; the value is in the order of magnitude.

How conversion gets gamed or misread. Changing the denominator — switching from "visitors" to "qualified visitors" improves the rate without improving the business. Measuring conversion at day 30 when your trial is 14 days, capturing people who converted for reasons unrelated to the trial. Reporting conversion on a mixed cohort of traffic sources, so a shift in traffic mix looks like a product improvement. Counting a $1 introductory offer as a conversion. Ignoring involuntary churn — RevenueCat found billing failures caused 31% of Google Play cancellations and 14% of App Store cancellations, which is a payments problem misfiled as a product problem. And the structural one: optimizing top-of-funnel conversion by lowering intent, which raises the rate and lowers retained revenue.


A9. CAC payback period

Definition. The number of months of gross profit from a new customer required to repay the cost of acquiring them. This is the most reliable single measure of go-to-market efficiency, because it uses only realized data and cannot be extrapolated into fantasy the way LTV can.

Formula. CAC payback (months) = S&M expense (prior period) ÷ (New ARR added × Gross margin %) × 12

Two details matter enormously. Use gross-margin-adjusted revenue, not raw ARR — you repay acquisition cost out of gross profit, not revenue. And lag the S&M spend to the period that actually generated the bookings, typically one quarter.

Worked example. Q1 S&M spend was $600,000. Q2 new ARR added was $900,000. Gross margin is 75%. CAC payback = 600,000 ÷ (900,000 × 0.75) × 12 = 600,000 ÷ 675,000 × 12 = 10.7 months

If you had skipped the gross-margin adjustment, you would report 8.0 months — a 33% understatement. If you had used same-period spend in a quarter where you cut S&M, you would report something better still and entirely fictional.

What good looks like. Median CAC payback was 16 months in 2025, with the top quartile at 6 months or less and the bottom quartile at 24 months or more. By ACV, sub-$5K deals paid back in 11 months while $50K–$100K enterprise deals took 22 months. By growth rate, companies growing above 50% showed a 10-month median. Horizontal SaaS 14 months, vertical SaaS 18 months (Aleph × Benchmarkit 2026, n=198).

The practical reading: under 12 months is genuinely efficient, 12–18 months is normal, over 24 months means you are financing your customers' adoption with investor capital and need either a lot of capital or a very high retention rate to make it work. Note the direction of the ACV finding — enterprise deals are less efficient on payback and more efficient on retention. That trade is the whole strategic choice between motions.

How it gets gamed or misread. Omitting the gross margin adjustment (systematically understates payback by 20–35%). Using same-period rather than lagged spend. Using net new ARR (including expansion) against new-customer acquisition cost, which conflates two different efficiencies with very different costs — recall that expansion CAC ratio runs about half of new-customer CAC ratio. Excluding sales salaries. Computing payback on a blended base that mixes a fast-paying self-serve segment with a slow enterprise one.


A10. Contribution margin

Definition. Revenue from a unit minus all variable costs attributable to it — the money each incremental unit contributes toward fixed costs and profit. In software the unit is usually a customer or a cohort; in marketplaces an order; in hardware a device.

Formula. Contribution margin per unit = Unit revenue − Unit variable costs (including COGS, variable acquisition cost, transaction fees, support, refunds, chargebacks, and delivery or fulfilment where applicable)

Worked example — a delivery marketplace order. Order GMV $40. Take rate 22%, so gross revenue $8.80. Courier payout subsidy $2.10. Payment processing $1.46. Support and insurance allocation $0.35. Promotional discount amortized per order $1.20. Contribution margin = 8.80 − 2.10 − 1.46 − 0.35 − 1.20 = $3.69 per order, or 42% of net revenue and 9.2% of GMV.

Now suppose acquisition cost is $28 per new customer and a customer orders 1.4 times per month. Months to recover acquisition cost = 28 ÷ (3.69 × 1.4) = 5.4 months. The whole business hangs on whether customers stay past month six.

What good looks like. Contribution margin has no universal benchmark because the unit definition varies. The useful test is directional: contribution margin must be positive and must improve with scale. If it is negative, growth destroys cash and every incremental customer makes things worse — this is the "negative unit economics" failure mode, and it is survivable only if there is a specific, dated mechanism by which it turns positive (a take-rate increase, a fulfilment density effect, a cost curve). "We'll fix it with scale" without a named mechanism is not a plan.

How it gets gamed or misread. Excluding promotional discounts and credits from variable costs (they are the most variable cost you have). Excluding refunds and chargebacks. Treating support as fixed when it scales with customers. Presenting "contribution margin after variable costs but before customer acquisition" as if acquisition were optional. Computing it on the best cohort. And the subtle one: presenting steady-state contribution margin — what a mature customer generates — as if it applied to all customers, when early-life customers cost more to support and churn more.


A11. Sales efficiency: magic number, burn multiple, Rule of 40

These three answer the same question — how much are you spending to buy growth? — from different angles.

Magic number

Definition. New ARR generated per dollar of prior-period sales and marketing spend. Originated in public SaaS analysis and usually computed quarterly.

Formula. Magic number = New ARR added in period ÷ S&M expense in prior period

Worked example. Q3 S&M spend $1,200,000. Q4 new ARR added $1,500,000. Magic number = 1,500,000 ÷ 1,200,000 = 1.25.

What good looks like. The conventional thresholds: below 0.75 means the go-to-market motion is not paying for itself and you should fix efficiency before adding spend; 0.75–1.0 is acceptable; above 1.0 suggests you should spend more. The 2025 actuals: median 1.37, up sharply from 0.94 in 2024; 75th percentile 2.14; bottom quartile 0.68. Companies growing above 50% showed a median of 2.40, while those growing 11–30% sat below 0.75 (Aleph × Benchmarkit 2026, n=132).

That year-over-year jump from 0.94 to 1.37 is one of the most significant findings in the 2026 benchmark cycle, and it reflects the post-2022 efficiency reset more than any improvement in demand.

How it gets gamed. Not lagging the spend. The Aleph guidance names this explicitly as the most common error: dividing current-period new ARR by current-period spend "flatters a ramping business and penalizes one that just cut." A company that slashes S&M in Q4 shows a spectacular Q4 magic number and a terrible Q2. Read it as a rolling four-quarter trend.

Burn multiple

Definition. David Sacks's formulation: how much cash the company burns to generate each incremental dollar of ARR. It is a whole-company efficiency metric, not a sales metric — it catches waste anywhere in the business, which is exactly why it is useful.

Formula. Burn multiple = Net burn ÷ Net new ARR

Sacks introduced it as "an annualized version of the Hype Ratio," arguing that it "puts the focus squarely on burn by evaluating it as a multiple of revenue growth," and that unlike cumulative-capital metrics it "doesn't care about sunk costs and always gives founders the chance to improve" (Craft Ventures / David Sacks, April 23, 2020).

Sacks's benchmark table for venture-stage startups:

Burn multiple Rating
< 0.5x Amazing
0.5x – 1x Great
1x – 2x Good
2x – 3x Suspect
> 3x Bad

Worked example. A company burns $4,000,000 net over a year and grows ARR from $6,000,000 to $9,500,000. Net new ARR = $3,500,000. Burn multiple = 4,000,000 ÷ 3,500,000 = 1.14x — "good" on the Sacks scale.

Now suppose it burned the same $4,000,000 but ARR went from $6,000,000 to $7,200,000. Burn multiple = 4,000,000 ÷ 1,200,000 = 3.33x — "bad." Same spend, very different company.

What good looks like in practice. Burn multiple should fall as a company scales. Benchmarkit's 2025 report indicates the target compresses below 1.0x in the $25M–$50M ARR range (Benchmarkit 2025). Practitioner guidance for 2026 places seed-stage burn multiples around 2.5–3.4x, Series A medians near 1.6x, growth-stage targets near 1.4x, and scale-stage targets at or below 1.0x (Runway, burn multiple benchmarks 2026) — note that these are practitioner estimates rather than a documented survey panel, and should be treated as directional.

How it gets gamed or misread. Using gross new ARR instead of net new ARR, which ignores churn and is the most common manipulation. Using a quarter with a large annual prepayment, which temporarily suppresses net burn. Excluding capex or capitalized software from burn. The metric is also genuinely unhelpful at very early stage: when net new ARR is near zero, the ratio explodes to infinity and tells you nothing you did not already know.

Rule of 40

Definition. Growth rate plus profit margin should exceed 40. It is a maturity metric — a way of saying that a company should be trading growth against profitability along a defined frontier.

Formula. Rule of 40 = YoY revenue growth % + EBITDA margin % (free cash flow margin or operating margin also acceptable, applied consistently)

Worked example. A company grows 32% with an EBITDA margin of −6%. Score = 26. A company grows 12% with a 20% margin. Score = 32. The second is "better" on this metric despite growing a third as fast.

What good looks like. Median Rule of 40 in 2025 was 25, up 10 points from 15 in 2024; top quartile 43, bottom quartile 7. The metric is only meaningful at roughly $20M ARR and above — below that, growth rate and unit economics are the more reliable signals (Aleph × Benchmarkit 2026, n=110).

How it gets misread. Applying it to a $2M ARR company, where it is noise. Using "adjusted EBITDA" with stock compensation, restructuring and "one-time" items excluded, which can add 20 points. Treating 40 as a pass/fail line rather than a frontier — a company at 35 growing 80% is in a completely different position from one at 35 growing 5%.

ARR per employee

A useful cross-check on all of the above. Median ARR per employee in 2025 was $193,000, up 29% from $150,000 in 2024, with top quartile ~$279,000 and bottom quartile $126,000; peak efficiency occurs in the $20–50M ARR band at $282,000 (Aleph × Benchmarkit 2026, n=96). SaaS Capital's broader and smaller-company-weighted panel reports a lower median of $141,125, up from $129,724 the prior year, with bootstrapped companies consistently higher than equity-backed at every ARR level ($177,240 vs $152,295 in the $5–10M band) (SaaS Capital, published July 30, 2026).

The 29% single-year jump in the Benchmarkit panel is worth pausing on: it reflects both revenue growth and deliberate headcount reduction, and is the clearest quantitative trace of AI-assisted productivity plus post-2022 discipline showing up in operating metrics. The two panels disagree by 37%, which is a good reminder of how much panel composition drives benchmark figures.


A12. Product engagement: DAU/MAU, activation, quick ratio

DAU/MAU (stickiness)

Definition. The share of monthly active users who use the product on a given day. A proxy for how many days per month a typical user shows up.

Formula. DAU/MAU = average daily active users ÷ monthly active users

Worked example. 40,000 MAU, 9,200 average DAU. DAU/MAU = 23%, implying the average monthly user shows up about 7 days a month.

What good looks like. Entirely dependent on the product's natural frequency. A messaging app should be 50%+; a social feed 40–60%; a productivity tool used on workdays tops out near 70% even when beloved (5 of 7 days); a tax product or an HR system might be 3% and perfectly healthy. The benchmark is not a number, it is your own product's theoretical maximum. Ask: if a user loved this as much as possible, how often would they open it? Measure against that.

How it gets gamed or misread. Counting a push notification open as an active session. Defining "active" as any app launch rather than a meaningful action. Comparing DAU/MAU across products with different natural frequencies — the most common analytical error in consumer product reviews. Using it as a north star, which incentivizes engagement-farming mechanics (notification spam, streaks, artificial scarcity) that raise the ratio and lower long-run retention.

Activation rate

Definition. The share of new users who reach a defined first-value moment within a defined window.

Formula. Activation rate = new users reaching the activation event within N days ÷ new users in cohort

Worked example. Activation is defined as "created a project and invited one teammate within 7 days." Of 1,400 signups in March, 392 did so. Activation = 28%.

What good looks like. Median 25% overall, 30% for SaaS, per Lenny Rachitsky's 500+-response survey (October 25, 2022). The range is enormous: B2C freemium products with easy activation milestones run highest; marketplaces and e-commerce, where activation means a first transaction, run lowest.

How it gets gamed. Redefining the activation event to something easier. This is the purest example of Goodhart's law in product analytics: when activation becomes a target, teams move the definition rather than the behavior. Lock the definition, and if you must change it, restate history.

Quick ratio

Definition. Growth efficiency at the revenue level — how much revenue you add for every dollar you lose. Popularized by Social Capital's analysis of SaaS growth.

Formula. Quick ratio = (New MRR + Expansion MRR) ÷ (Churned MRR + Contraction MRR)

Worked example. New $6,000, expansion $2,000, churned $3,000, contraction $1,500. Quick ratio = 8,000 ÷ 4,500 = 1.78.

What good looks like. The widely-used threshold is 4 or above for an efficiently growing SaaS business; below 1 means the business is shrinking. A quick ratio between 1 and 2 means you are running very hard to move slowly — a company with a quick ratio of 1.2 growing 20% is adding six dollars of new revenue for every five it loses, and almost all of its sales effort is replacement.

How it gets misread. It is scale-dependent: a company at $50,000 MRR can post a quick ratio of 10 trivially, while one at $50M cannot. It also hides absolute magnitude — 4.0 on tiny numbers is meaningless. And it says nothing about cost; a company buying growth at a terrible CAC can have a fine quick ratio.


A13. Marketplace metrics: GMV, take rate, liquidity

GMV (gross merchandise value)

Definition. Total value of transactions flowing through the marketplace. GMV is not revenue — a point worth repeating because the conflation is endemic (see A1.3).

Formula. GMV = Σ (transaction value of all completed orders in period)

Worked example. 42,000 orders at an average of $58. GMV = $2,436,000. At an 18% take rate, revenue = $438,480. These are numbers that differ by 5.6x, and only one of them is yours.

Take rate

Definition. The share of GMV the marketplace keeps as revenue.

Formula. Take rate = Net revenue ÷ GMV × 100

What good looks like. CJ Gustafson's framework is the most useful available: take rate is a function of purchase frequency, ticket size and platform labor intensity, producing roughly three tiers — ~10% for lead generation (Amazon, eBay, Thumbtack: the platform introduces the parties and steps back), 10–20% for trust-layer marketplaces (Airbnb ~15% on the supply side, Etsy, StockX: the platform provides payments, reviews, dispute resolution), and 20–30% for heavily managed logistics (Uber, Lyft, DoorDash, Vacasa: the platform operates the service) (Mostly Metrics, January 18, 2026).

The strategic implication is that take rate is not a pricing decision you get to make freely — it is set by how much work you do. Raising take rate without adding value invites disintermediation.

Liquidity and match rate

Definition. Liquidity is the probability that a listing or a search results in a transaction within a reasonable time. It is the core health metric of a marketplace and the one most often skipped in favor of GMV.

Formulas (pick the one that matches your marketplace):

  • Supply-side liquidity = listings resulting in a transaction within N days ÷ total listings
  • Demand-side liquidity (match rate) = searches/requests resulting in a transaction ÷ total searches/requests
  • Time to match = median hours or days from listing/request to transaction

Worked example. In a freelance marketplace, 3,200 jobs are posted in a month and 1,376 are filled within 14 days. Supply-side liquidity = 43%. Median time to first qualified applicant: 9 hours. Median time to hire: 4.2 days.

What good looks like. There is no cross-category benchmark, because liquidity thresholds depend on what the buyer's alternative is. The rules that hold: liquidity must be measured within a geography and category, not globally — a rideshare marketplace with 70% national match rate and 20% in your city is unusable in your city. And liquidity is the leading indicator; GMV is the lagging one. Liquidity deteriorating while GMV grows means you are buying transactions with subsidy.

GMV retention

Definition. Cohort-based measure of how much of a cohort's spend (demand side) or earnings (supply side) you retain over time. Olivia Moore's framing at a16z: best-in-class supply-side GMV retention sits at or above 100% through month 12 and beyond, with a typical pattern of 80–95% in the first three months plateauing around 45–50% by month 12; in best-in-class marketplaces, 50–70% of suppliers retain but expand their GMV 2–3x (a16z, April 28, 2022). Demand-side retention is generally lower than supply-side.

How marketplace metrics get gamed or misread. Reporting GMV as revenue — still the most common. Reporting gross bookings including cancelled and refunded orders. Counting "annualized GMV run rate" off a holiday month. Reporting take rate before promotional spend, when discounts and courier subsidies come out of the same pocket. Measuring liquidity across the whole marketplace rather than per category-geography. And counting first-party inventory sales in "marketplace GMV."


A14. Utilization (for services and services-heavy businesses)

Definition. The share of available employee hours that are billable to clients. The central metric for agencies, consultancies, and the services arm of a software company.

Formula. Billable utilization = billable hours ÷ available hours (available = total capacity net of holiday and paid leave; "target utilization" is usually set well below 100% to allow for sales, admin and training)

Worked example. A consultant has 2,080 annual hours, minus 200 hours of holiday and leave = 1,880 available. They bill 1,250. Utilization = 66.5%. At a $185 blended rate, they generate $231,250 of revenue. If their fully-loaded cost is $175,000, gross margin on that person is 24%. Raising utilization to 75% — 1,410 hours — generates $260,850 and a 33% margin. Nine points of utilization is nine points of margin, and this is why services businesses are so operationally sensitive.

What good looks like. SPI Research's 2026 benchmark, covering 500+ professional services firms with 245,000+ employees and $63B of revenue (2025 data), found billable utilization of 66.4% — the lowest in the survey's history and well below the 75% target. Alongside it: project margin 37.7% (up from 35.9%), revenue per employee $168K, revenue per billable consultant $210K, attrition 11.3%, EBITDA 9.9% (Deltek / SPI Professional Services Maturity Benchmark 2026).

For a software company with a services arm, the relevant benchmark is different: professional services gross margin runs around 30% at the median (Benchmarkit 2025), which is why a company with 25% of revenue in services carries a blended gross margin roughly 12 points below its software margin.

How utilization gets gamed or misread. Counting internal projects as billable. Counting hours billed but written off before invoicing (the gap between utilization and realization is where agency profit disappears). Excluding non-billable staff from the denominator so a firm with heavy overhead looks efficient. Pushing utilization above 85%, which reliably produces burnout and attrition — and at 11% attrition, replacing a consultant costs far more than the marginal billable hours gained. Target utilization is an optimization with an interior maximum, not a number to maximize.


A15. Unit economics, overall

Definition. Whether a single unit of your business — one customer, one order, one device — makes money after all the costs required to get and serve it. Everything in this chapter is ultimately a component of this question.

The general form. Unit contribution = Unit revenue − Unit COGS − Unit variable operating cost − Unit acquisition cost

Plus a time dimension: how many months until cumulative unit contribution exceeds acquisition cost, and what fraction of units survive that long.

Worked example, done properly, for a B2B SaaS customer.

Line Value
ACV $9,600
Gross margin 76%
Annual gross profit $7,296
Fully-loaded CAC $9,100
CAC payback 15.0 months
Annual logo churn 14%
Implied median customer life ~4.6 years
3-year cumulative gross profit (with 14% annual churn, undiscounted) ~$18,900
3-year contribution after CAC ~$9,800
3-year return on CAC 2.08x

This customer is profitable — but the company recovers nothing for 15 months, and 14% of customers do not survive the first year. At 22% annual churn instead of 14%, three-year cumulative gross profit falls to about $16,700 and return on CAC to 1.8x. At 22% churn and an 18-month payback, the business needs a lot of patient capital.

The three questions that matter.

  1. Is unit contribution positive today? If not, name the specific dated mechanism that makes it positive.
  2. How long until a unit repays its acquisition cost, and do enough units survive that long?
  3. Does unit economics improve or degrade with scale? Improving: density effects, fixed-cost leverage, pricing power, negotiated input costs. Degrading: rising CAC as you exhaust the best channels, worse-fitting customers, higher support load, higher inference load per user as engagement deepens.

How unit economics gets gamed or misread. Using best-cohort data. Excluding acquisition cost ("unit economics excluding CAC" is a phrase that should end a meeting). Using steady-state costs rather than actual current costs. Assuming fixed costs stay fixed. And the systemic error: modelling unit economics at current scale and applying them at 10x scale, where CAC is higher, customer quality is lower, and support load is heavier.


A16. Benchmark summary (2025–2026 data)

All figures are medians on surviving, mostly funded, self-reporting companies unless noted. Read them as "the middle of the companies still standing," not "what to expect."

Metric Median Top quartile Source / panel
YoY growth, private B2B SaaS 22% SaaS Capital 2026, 1,000+ cos, 2025 data
YoY growth, bootstrapped 20% 42.3% (90th pct, $3–20M ARR) SaaS Capital 2026
YoY growth, equity-backed 25% SaaS Capital 2026
Software gross margin 80% 86%+ Aleph × Benchmarkit 2026, n≈230
Total revenue gross margin 76% Aleph × Benchmarkit 2026
Gross margin, usage-only pricing 62% Aleph × Benchmarkit 2026
Net revenue retention 102% 110% Aleph × Benchmarkit 2026, n=230
NRR, usage-based / seat-based 108% / 98% 155% (usage, 75th pct) Aleph × Benchmarkit 2026
Gross revenue retention 84% (was 88%) 91% Aleph × Benchmarkit 2026, n=226
CAC payback 16 months ≤6 months Aleph × Benchmarkit 2026, n=198
Magic number 1.37 (was 0.94) 2.14 Aleph × Benchmarkit 2026, n=132
Rule of 40 25 (was 15) 43 Aleph × Benchmarkit 2026, n=110
ARR per employee $193K $279K Aleph × Benchmarkit 2026, n=96
ARR per employee (broader panel) $141K SaaS Capital 2026, 1,000+ cos
Consumer sub 12-month retention ~27–28% RevenueCat 2026, 115K+ apps
Consumer trial→paid, hard paywall 10.7% (day 35) RevenueCat 2026
Activation rate, SaaS 30% Lenny's Newsletter 2022, n=500+
Billable utilization, services 66.4% 75% (target) SPI/Deltek 2026, 500+ firms

PART B — EIGHT ILLUSTRATIVE FINANCIAL MODELS

How to read these models

These are illustrative arithmetic, not forecasts, not templates, and not advice. Each is built from explicitly stated assumptions anchored where possible to the cited 2025–2026 benchmark data in Part A. Their purpose is to show how the metrics interact, where cash actually runs out, and what the gap looks like between the version in a pitch deck and the version that usually happens.

Each model shows two paths:

  • "If it works" — roughly a 75th–85th percentile outcome conditional on the company surviving. This is the shape of a pitch deck.
  • "Typical" — roughly a median outcome conditional on the company surviving long enough to have numbers at all.

Both paths condition on survival, which is the largest single distortion in this entire chapter. The unconditional median outcome for a newly formed startup is not the "typical" column — it is closure, acquihire, or indefinite stagnation at a revenue level too small to matter. Read "typical" as "typical among the ones that made it three years."

All figures are rounded. Taxes, working-capital timing, and equity accounting are simplified or omitted.


Model 1 — Bootstrapped micro-SaaS

Shape. One founder, one part-time contractor. A B2B utility product at $49/month. Distribution by content, SEO and community — no paid acquisition, because there is no capital to deploy. 36-month horizon.

Assumptions

Assumption Value Basis
Price (ARPA) $49/month
Gross margin 85% Above the 80% software median; no support org, minimal infra (Aleph × Benchmarkit 2026)
Monthly logo churn 3.5% (≈35% annual) Worse than the 80% GRR benchmark for sub-$5K ACV, because that panel samples larger survivors
Paid CAC $0 Content/SEO only
Founder opportunity cost $8,000/month Not in the P&L; the real cost of this model
Fixed cash costs $900/month rising to $2,400 Hosting, tools, contractor
Starting capital $15,000 of savings

Trajectory (MRR)

Month Typical: customers Typical: MRR If it works: customers If it works: MRR
6 29 $1,400 54 $3,200
12 79 $3,900 153 $9,000
18 124 $6,100 288 $17,000
24 161 $7,900 455 $26,800
30 190 $9,300 615 $36,300
36 214 $10,500 748 $44,200

(If-it-works path assumes 3.0% monthly churn and a price increase to $59 at month 10.)

Key ratios at month 36

Typical If it works
ARR $126,000 $530,000
Gross margin 85% 86%
Net margin (pre-founder-salary) 71% 78%
Burn multiple n/a — never burned n/a
CAC payback ~0 months (no paid CAC) ~0 months
Steady-state ceiling (adds ÷ churn) ~314 customers, $15,400 MRR ~1,300 customers

Where it breaks. Three places, in order of likelihood.

  1. The churn ceiling. With 3.5% monthly churn and 11 new customers a month, the business asymptotes at 314 customers (11 ÷ 0.035) — about $15,400 MRR. The ceiling is set by the ratio of acquisition rate to churn rate, and no amount of time raises it. A founder grinding toward month 48 at constant effort does not get to $30,000 MRR; they get to $15,000 and stay there. Breaking the ceiling requires raising the numerator (more acquisition, which means paid channels and therefore capital) or lowering the denominator (better retention, which usually means moving upmarket to larger customers).
  2. Founder capacity. At 214 customers, support load is roughly 15–20 hours a month. At 750 it is a full-time job that the founder is also trying to do alongside product and marketing. The typical failure is not bankruptcy; it is the founder taking a job.
  3. Channel decay. An SEO-led micro-SaaS is one search-algorithm change from losing most of its acquisition. Content-driven acquisition has become materially more volatile since AI-generated answers began absorbing informational search traffic.

Honest framing. The typical column here produces about $89,000 of annual gross profit against a founder opportunity cost of roughly $96,000. On a pure cash basis, this is a job that pays slightly less than a job — plus an asset that might sell for 3–5x ARR. That can be an excellent outcome if the work is 15 hours a week, and a poor one if it is 50. The unconditional median micro-SaaS never reaches $1,000 MRR.


Model 2 — Venture-backed B2B SaaS

Shape. Post-seed company, $3.5M raised, $600K ARR at month 0, 14 people, mid-market product at $24,000 ACV. 36-month horizon covering the Series A decision.

Assumptions

Assumption Value Basis
Starting ARR $600,000 (25 customers @ $24K)
Gross margin 76% (typical) / 76% (works) Total-revenue median (Aleph × Benchmarkit 2026)
Annual GRR 88% (works) / 82% (typical) 84% panel median, 88% for sales-led motion
NRR 112% (works) / 98% (typical) 102% panel median; 110% top quartile
CAC payback 14 months (works) / 22 months (typical) 16-month median; $50K–$100K ACV cohort at 22
Starting opex $190,000/month 14 FTE plus overhead
Starting cash $3,500,000

Trajectory

Month Typical ARR Typical opex/mo Typical net burn/mo Typical cash Works ARR Works net burn/mo Works cash
6 $813,000 $215,000 $164,000 $2,555,000 $979,000 $165,000 $2,540,000
12 $1,102,000 $249,000 $181,000 $1,511,000 $1,597,000 $170,000 $1,529,000
18 $1,347,000 $267,000 $184,000 $413,000 $2,291,000 $178,000 $14,477,000*
24 $1,647,000 $287,000 $186,000 $2,302,000** $3,287,000 $178,000 $13,402,000
30 $1,898,000 $294,000 $177,000 $1,218,000 $4,281,000 $177,000 $12,335,000
36 $2,189,000 $301,000 $166,000 $193,000 $5,574,000 $166,000 $11,305,000

* $14M Series A closes at month 18. ** $3M insider bridge at month 20.

Key ratios

Typical If it works
Year-2 growth 49% 106%
Year-3 growth 33% 70%
Burn multiple, year 2 4.1x ("bad") 1.26x ("good")
Burn multiple, year 3 3.9x 0.92x
Rule of 40 at month 36 33 + (−91) = −58 70 + (−36) = 34
ARR per employee at month 36 ~$128,000 ~$175,000

Where it breaks. At month 18 in the typical column, with $413,000 of cash and a $184,000 monthly burn — about 2.2 months of runway, which is a crisis, not a fundraise. The company had to start raising at month 12 with $1.1M ARR and 49% growth, which in 2026 is not a Series A profile. The realistic path is a bridge from existing investors at flat or modest terms, which buys 12 months and lands the company at month 36 with $2.2M ARR, no cash, and a burn multiple that has been above 3x for two years.

The mechanical cause is visible in the table: opex grew every month while ARR growth decelerated. Burn stayed roughly constant in dollars while net new ARR shrank, which is precisely what a rising burn multiple measures.

The lesson the numbers teach. The two columns start identically and diverge from a growth-rate difference of about 3 percentage points per month. Compounded over 36 months that is a 2.5x difference in ending ARR and the difference between a $14M Series A and a survival bridge. Small, sustained differences in growth rate dominate every other variable in a venture-backed model — which is exactly why founders are tempted to buy growth at any CAC, and exactly why burn multiple exists as a check on that temptation.


Model 3 — Consumer app (advertising + subscription)

Shape. A mobile app with a free ad-supported tier and a $9.99/month or $59.99/year subscription behind a hard paywall. 24-month horizon.

Assumptions

Assumption Typical If it works Basis
Paid CPI $3.20 $3.20 Market rate, Tier-1
Organic share of installs 45% 55%
Install → paying subscriber 3.0% 5.5% Hard-paywall day-35 trial→paid median 10.7%; install→trial ~25–30% (RevenueCat 2026)
App store fee 15% 15% Small-business program rate
12-month subscriber retention 28% 47% Panel median ~27–28%
Ad ARPDAU (free users) $0.012 $0.016
DAU/MAU 18% 22%

Unit economics per install, 12-month horizon

Line Typical If it works
12-month revenue per paying subscriber (net of store fee) $38.24 $63.35
Subscription revenue per install $1.15 $3.48
Ad revenue per install $0.23 $0.43
Total 12-month revenue per install $1.37 $3.92
vs. paid CPI of $3.20 0.43x — loses money 1.22x — marginal
vs. blended cost per install (incl. organic) 0.78x — still loses money 2.72x — works

Trajectory (typical path)

Quarter Installs/mo Paying subs MRR Ad revenue/mo Net monthly result
Q1 40,000 1,100 $9,300 $2,100 −$62,000
Q2 75,000 2,900 $24,600 $4,300 −$88,000
Q3 120,000 5,200 $44,200 $7,600 −$121,000
Q4 160,000 7,600 $64,600 $11,200 −$148,000
Q6 (M18) 190,000 10,100 $85,900 $15,400 −$171,000
Q8 (M24) 200,000 11,400 $96,900 $17,900 −$180,000

(Assumes paid UA of $185K/month at Q3 onward, plus $95K/month of fixed opex.)

Where it breaks. Immediately and structurally: paid user acquisition is unprofitable on a 12-month horizon in the typical column, and scaling it accelerates cash loss. Every incremental dollar of UA spend returns $0.43 within a year. The only paths out are (a) organic amplification large enough to drag blended cost below revenue per install, (b) a step change in conversion or retention, or (c) stopping paid acquisition entirely and accepting organic-only growth.

The failure mode this produces is distinctive: the app looks like it is working right up until the moment UA spend stops. MRR grows every month, install counts grow every month, and the deck shows a beautiful curve. The moment the UA budget is cut, MRR begins decaying at the churn rate — and with 72% of annual subscriptions cancelling within the first year, the decay is fast.

Key ratios

Typical If it works
Revenue per install ÷ paid CPI 0.43x 1.22x
Payback on paid install Never (within 12 months) ~11 months
Month-12 subscriber retention 28% 47%
Share of revenue from ads 16% 11%

A note on the ad tier. In this model advertising contributes 11–16% of revenue while consuming most of the user base and most of the support load. That is characteristic. Advertising monetizes at roughly one to two cents per daily active user in Tier-1 markets; a subscription monetizes at roughly $0.28 per subscriber-day. Ad revenue in a subscription app is usually a rounding error dressed up as a second business model. It is worth having as a floor and as a way to monetize non-converters, not as a strategy.


Model 4 — AI startup with usage-based inference costs

This is the central new economics question, so this model gets the most detail.

Shape. An AI agent that resolves customer-support tickets. Two pricing options are modelled — per-outcome ($0.60 per resolved ticket) and per-seat ($80/user/month) — because the choice between them determines whether the business has gross margins at all.

Assumptions

Assumption Value Basis
Tokens per resolution ~25,000 input + 4,000 output, across 4 model calls (agent loop with retrieval, reasoning, drafting, verification) Typical agentic workload
Frontier model price $5.00 / $25.00 per million input/output tokens BenchLM, September 2026
Production-grade small model price $0.75 / $3.75 per million Same
Other COGS per resolution $0.05 Vector DB, orchestration, logging, observability
Price per resolved ticket $0.60 Priced against ~$2.40 human cost per ticket

The gross-margin compression, shown explicitly

Phase Model routing Inference cost per resolution Gross margin at $0.60 price
Launch (months 1–6) 100% frontier model, no caching, full agent loop every time $0.900 −58%
Mid (months 7–14) 60% frontier / 40% small model $0.594 −7%
Optimized (months 15+) 20% frontier / 80% small model, prompt caching cuts input tokens 60%, loop reduced from 4 calls to 3 $0.234 +53%

This is the single most important table in this chapter. An AI product can launch at negative gross margin, look like a runaway success on revenue, and only reach software-like economics after 12–18 months of deliberate cost engineering. The engineering is real work — routing logic, evaluation harnesses to prove the cheap model is good enough, caching, context pruning, loop reduction — and it competes directly with feature work for the same engineers.

Why seat pricing is dangerous for a usage-driven product

Same product, priced at $80/user/month plus $4/user of other COGS:

Usage intensity Launch-phase GM Mid-phase GM Optimized GM
Light user — 150 resolutions/month −74% −16% +51%
Medium user — 400 resolutions/month −355% −202% −22%
Heavy user — 900 resolutions/month −918% −573% −168%

Under seat pricing, your best customers destroy you. The heaviest users — the ones with the deepest adoption, the ones your case studies feature, the ones who would never churn — generate the largest losses. Every retention and expansion incentive in the business is pointed at the thing that kills it. This is the specific, mechanical reason so many 2024–2026 AI companies shifted to usage-based, hybrid or outcome-based pricing, and it is not a preference; it is arithmetic.

The mitigations a16z describes — moving off unlimited introductory plans and rate-limiting the unprofitable top few percent of users (a16z, August 21, 2025) — are real, and this table shows why they are necessary rather than optional.

Trajectory

Month Typical: ARR Typical: GM Typical: burn/mo If it works: ARR Works GM Works burn/mo
6 $420,000 −40% $310,000 $1,400,000 −25% $420,000
12 $1,100,000 −5% $380,000 $4,800,000 +12% $560,000
18 $1,900,000 +28% $390,000 $11,000,000 +44% $520,000
24 $2,600,000 +44% $340,000 $19,000,000 +58% $310,000

Key ratios at month 24

Typical If it works
Gross margin 44% 58%
Inference as % of revenue 48% 34%
Inference efficiency ratio 2.1:1 2.9:1
Burn multiple, year 2 2.7x ("suspect") 0.35x ("amazing")
NRR 104% 148%

Against Ben Murray's suggested targets of 5:1 for AI-native products, both columns are short; the typical column at 2.1:1 is a business where roughly half of every revenue dollar goes to the model provider (The SaaS CFO, May 2026).

Where it breaks. Four distinct failure points, all of which have been observed in the market:

  1. Cost engineering doesn't land. If evaluation shows the cheap model fails on 12% of tickets and the customer's contract has an accuracy SLA, you cannot route away from the frontier model, and you are stuck at the mid-phase margin forever.
  2. The frontier ratchet. Token prices fall, but customers migrate to newer, more capable, more expensive models and to reasoning modes that burn 5–20x the tokens. Cost per token falls; cost per task rises. Gustafson's framing — old cheap models are "as desirable as a flip phone at an iPhone launch" — is the core objection to the "costs will fall" argument (Mostly Metrics, September 2025).
  3. Pilot revenue does not renew. A large share of early AI revenue is funded from innovation budgets. It annualizes into an impressive ARR figure and then does not renew. Consumer AI apps show a version of this: RevenueCat found AI apps retaining substantially worse at 12 months than traditional apps, even while showing higher realized LTV per subscriber ($30.16 vs $21.37) — high-value, short-lived customers.
  4. Growth outruns margin repair. In the "if it works" column, ARR grows 13x in 18 months while gross margin is climbing from negative territory. If growth is fast enough and margin repair is slow enough, the absolute dollar cost of negative-margin revenue can exceed the company's entire capital base. Growing fast at negative gross margin is the only situation in startups where success and failure are the same event.

How to read AI company metrics, given all this. Ask for: committed versus consumption revenue; gross margin with inference separately disclosed; trailing-twelve-month revenue, not annualized-month ARR; NRR on a fixed cohort; and the share of revenue from customers past their first renewal. Bessemer's 2025 segmentation is a useful reminder of the range — their fastest-growing "Supernova" cohort averaged roughly 25% gross margin, often negative, while their "Shooting Stars" cohort averaged 60% (Bessemer, August 2025). Both are AI companies. They are not the same kind of asset.


Model 5 — Hardware startup

Shape. A consumer hardware device sold direct-to-consumer at $249. 30-month horizon from tooling commitment through three production runs.

Assumptions

Assumption Value Basis
BOM $80/unit at 1,000-unit volume
Landed cost multiplier 1.30x BOM Freight, duty, inspection (Glencoyne, September 2025)
Tariff 15% on landed Within the cited 10–25% range
Fulfilment per unit $14 Pick, pack, ship, packaging
Return rate 6% Consumer electronics DTC
NRE / tooling $40,000 Molds, jigs, fixtures; simple molds from $5,000, complex multi-cavity $100,000+
Contract manufacturer deposit 40% of run value Cited range 30–50%
MOQ 1,000 units Cited typical range 1,000–5,000
Component lead time 20–50 weeks Cited
Cash cycle (deposit → customer payment) 8+ months Cited
DTC CAC $55/unit

Unit economics

Line Per unit
Retail price $249
Landed cost (BOM × 1.30) $104
Tariff (15%) $16
Fulfilment $14
Subtotal COGS $134
Returns-adjusted COGS (÷ 0.94) $142
Gross margin 43%
CAC $55
Contribution per unit $52 (21% of price)

The working-capital trap, quantified

Event Cash out Cumulative
Tooling / NRE $40,000 $40,000
40% deposit on 1,000 units $41,600 $81,600
Balance on shipment (60%) $62,400 $144,000
Freight + duty $16,000 $160,000
Certification (final) $28,000 $188,000
Cash out before first customer dollar $188,000
Revenue if all 1,000 units sell at $249 $249,000
Gross profit if all units sell ~$107,000

Where it breaks. The number that kills hardware companies is not gross margin; it is the eight-month gap between paying for inventory and being paid for it. To sell 1,000 units per quarter, you must fund the next quarter's inventory before the current quarter's revenue arrives. Growth consumes cash at roughly (units growth) × (landed cost per unit), permanently.

Scenario Quarterly units Inventory pre-funding required Quarterly gross profit
Typical 400 → 700 → 900 $73,000 → $94,000 $43,000 → $96,000
If it works 1,000 → 2,500 → 5,000 $260,000 → $520,000 $107,000 → $535,000

In the "if it works" column, faster growth makes the cash position worse before it makes it better. At 5,000 units a quarter the company is pre-funding $520,000 of inventory against $535,000 of quarterly gross profit — it is running at roughly breakeven on cash while doubling revenue, which is the classic hardware squeeze that pushes companies into inventory financing, revenue-based facilities, or a dilutive round raised from a position of weakness.

Other break points: the second production run. Run one is often funded by crowdfunding or pre-orders, where customers pay up front and working capital is free. Run two has no pre-orders. Many hardware companies die between run one and run two. And the gross margin is too thin for retail. At 43% direct margin, a retail channel taking 40–50% leaves nothing. A hardware company that wants retail distribution needs to design for a 60%+ direct margin from the start, which usually means a BOM under 20% of retail price.

Honest framing. A $249 device with an $80 BOM — a roughly 3.1x BOM-to-retail multiple — is on the thin side of viable. The conventional guidance of a 4–5x multiple exists precisely because of the cost stack above: tariffs, freight, returns, fulfilment and channel margin consume the difference between the naive margin (68%) and the real one (43%).


Model 6 — Marketplace

Shape. A local services marketplace (home repair). 30-month horizon covering liquidity build in a first city and expansion to three more.

Assumptions

Assumption Value Basis
Average order value $118
Take rate 15% rising to 18% Trust-layer tier: 10–20% (Mostly Metrics, January 2026)
Payment processing 2.9% + $0.30 per order
Support / trust & safety $0.40 per order
Supply-side incentives 3.0% of GMV falling to 0.5% Liquidity bootstrapping
Buyer CAC $28
Supplier CAC $190
Target supply-side liquidity 45% of requests filled in 48h

Trajectory

Period GMV Take rate Net revenue Orders Liquidity Contribution margin
Q2 (city 1) $120,000 15% $18,000 1,017 22% −$3,800
Q4 (city 1) $600,000 16% $96,000 5,085 38% $31,000
Q6 (cities 1–2) $1,800,000 17% $306,000 15,254 44% $164,000
Q8 (cities 1–4) $3,600,000 18% $648,000 30,508 41% $352,000

Key ratios at Q8

Typical If it works
GMV run rate $14.4M $38M
Net revenue $2.6M $6.8M
Take rate 18% 18%
Contribution margin (% of net revenue) 54% 61%
Buyer payback 4.9 months 2.8 months
Supply-side GMV retention, m12 41% 78%
Demand-side repeat rate, 12-month 31% 52%

Against the a16z benchmark — best-in-class supply-side GMV retention at or above 100% through month 12, with a typical plateau around 45–50% (a16z, April 2022) — the typical column at 41% is slightly below par and the "works" column at 78% is strong but not best-in-class.

Where it breaks. Three specific places.

  1. Liquidity per city, not in aggregate. The Q8 row shows aggregate liquidity slipping from 44% to 41% while GMV doubles — the signature of expanding into new cities faster than each one reaches liquidity. A marketplace with 60% liquidity in city one and 15% in cities three and four reports 41% and is failing in two of four markets. Report liquidity per city-category or it is not a metric.
  2. Disintermediation. In a high-ticket, repeat-relationship services category, both sides have a strong incentive to transact off-platform after the first match. The take rate is only defensible while the platform provides ongoing value — scheduling, payments, guarantees, dispute resolution. Measure the share of second and third transactions that stay on platform; if it drops sharply, the business is a lead-generation service with a marketplace's cost structure.
  3. Take-rate ceiling. The model raises take rate from 15% to 18% over 30 months, which contributes roughly $430,000 of annualized revenue at Q8 — meaningful, and also the easiest number to get wrong. Take-rate increases that are not accompanied by added platform value produce supplier attrition with a lag of two to four quarters, which shows up as falling liquidity before it shows up as falling GMV.

The subsidy question. Supply incentives fall from 3.0% to 0.5% of GMV in this model. That assumption is doing enormous work: at Q8, holding incentives at 3.0% instead of 0.5% would consume $90,000 per quarter and cut contribution margin by a quarter. A marketplace that cannot reduce its subsidy as liquidity builds does not have a marketplace; it has a purchasing program. The test is empirical and cheap: cut incentives 20% in one city and watch liquidity.


Model 7 — Regulated startup (fintech / digital health)

Shape. A B2B product serving healthcare providers, handling protected health information, with a payments component requiring money transmission capability. 36-month horizon. The defining features are a compliance cost layer that exists before any revenue, and a sales cycle roughly twice that of unregulated B2B software.

Assumptions

Assumption Value Basis
HIPAA compliance, year-1 setup (seed stage, 11–50 staff) $25,000–$85,000 Accountable, March 2026
HIPAA ongoing annual $10,000–$35,000 Same
SOC 2 Type II, first year $40,000–$80,000 all-in Audit, tooling, pen test, staff time
Money transmission: partner-bank / agent model, year 1 $120,000–$200,000 Sponsor bank fees, BSA/AML program, compliance officer
Money transmission: direct nationwide licensing ~$250,000–$350,000 year 1; $225,000–$280,000+ ongoing Brico, July 2026
ACV $85,000 Mid-size provider group
Sales cycle 9–14 months vs. 3–6 for comparable unregulated B2B
Security review / procurement add-on +2–4 months

The compliance cost layer, year by year (typical path)

Year 1 Year 2 Year 3
HIPAA program $55,000 $22,000 $30,000
SOC 2 Type II $65,000 $45,000 $50,000
BSA/AML + sponsor bank (agent model) $160,000 $185,000 $210,000
Compliance headcount (0.5 → 1.5 FTE) $95,000 $180,000 $290,000
Outside counsel (regulatory) $70,000 $55,000 $80,000
Total compliance layer $445,000 $487,000 $660,000
Revenue $180,000 $850,000 $2,100,000
Compliance as % of revenue 247% 57% 31%

Trajectory

Month Typical ARR Customers If it works: ARR Customers
12 $180,000 2 $340,000 4
18 $425,000 5 $935,000 11
24 $850,000 10 $2,040,000 24
30 $1,445,000 17 $3,825,000 45
36 $2,100,000 25 $6,290,000 74

Key ratios at month 36

Typical If it works
Gross margin 64% (compliance ops partly in COGS) 71%
CAC payback 26 months 17 months
Annual GRR 93% 96%
NRR 108% 124%
Compliance as % of revenue 31% 12%
Burn multiple, year 3 2.9x 1.1x

Where it breaks. Two places, and they are related.

  1. The pre-revenue compliance bill. In year one the company spends $445,000 on compliance to earn $180,000 of revenue. That money is spent before product-market fit is proven, which means a regulated startup must raise more capital earlier than an unregulated one, on weaker evidence. Founders routinely underestimate this by 2–3x because they budget for the audit and not for the compliance officer, the outside counsel, the annual re-audit, or the twelve weeks of engineering time that security questionnaires consume.
  2. Sales cycle and cash collide. With a 9–14 month sales cycle plus 2–4 months of security review, a deal started in month 6 closes around month 20. The company must fund roughly 20 months of operation before the first cohort of real deals lands. In the typical column, CAC payback of 26 months means that even after closing, each customer takes over two years to repay acquisition cost.

The offsetting advantage, which is real. Look at the retention rows: 93% GRR and 108% NRR in the typical column, against a panel median of 84% GRR. Regulated products are extremely sticky — the switching cost includes re-running the buyer's own compliance process. Regulated startups trade a brutal first 24 months for durable revenue afterwards. The model only works if the company is capitalized to survive the trade.


Model 8 — Deep tech (grant-funded, long pre-revenue)

Shape. A materials-science company with a defensible process advantage. 48-month horizon (longer than the others, because 36 months does not reach revenue). Funding mixes non-dilutive grants with equity.

Assumptions

Assumption Value Basis
SBIR Phase I ~$305,000 (NSF); 2026 statutory cap $314,363 GrantProbe, May 2026
SBIR Phase I duration 6–12 months Same
SBIR Phase I award rate ~10–20%, ~14% historical average Same
Phase I → Phase II advancement ~40% at DoD, NIH, NSF Same
SBIR Phase II up to ~$2.0M over 24 months; 2026 cap $2,095,748 Same
Apply-to-funding lag 9–12 months Same
Pre-seed / seed equity $2.5M at month 6
Series A $12M at month 30, contingent on pilot data
Time to first commercial revenue Month 34 (typical) / month 26 (works)

Funding and burn trajectory (typical path)

Period Non-dilutive in Equity in Burn/mo Cash at period end Milestone
M1–6 $0 $2,500,000 $85,000 $1,990,000 Grant applications submitted
M7–12 $305,000 (Phase I) $0 $110,000 $1,635,000 Lab-scale validation
M13–18 $0 $0 $145,000 $765,000 Phase II application; gap period
M19–24 $900,000 (Phase II tranche 1) $1,500,000 (bridge) $190,000 $2,025,000 Pilot reactor build
M25–30 $700,000 (Phase II tranche 2) $0 $235,000 $1,315,000 Pilot data
M31–36 $400,000 $12,000,000 (Series A) $290,000 $11,975,000 First paid pilot, $250K
M37–48 $0 $0 $420,000 $5,955,000 $1.4M revenue, 3 customers

Key ratios

Typical If it works
Months to first revenue 34 26
Total capital consumed to first revenue ~$7.2M ~$5.8M
Non-dilutive share of capital raised 18% 24%
Revenue at month 48 $1.4M $6.5M
Gross margin at month 48 31% (pilot-scale) 48%
Burn multiple, year 4 n/a — meaningless pre-scale 2.6x

Note on burn multiple. For a company with near-zero revenue, net burn ÷ net new ARR divides by approximately zero and returns a meaningless number. Do not apply capital-efficiency ratios to pre-revenue deep tech. The right metrics are milestone-based: technical risk retired per dollar, time to next de-risking event, and whether each funding tranche buys a step change in the probability of the next one.

Where it breaks. The most dangerous point is months 13–18: the grant gap. Phase I ends, Phase II has not been awarded, the apply-to-funding lag is 9–12 months, and the equity raised at month 6 is running out. This gap is structural — it exists because grant cycles and company cash cycles are not synchronized — and it is where a large share of grant-funded deep tech companies fail.

Compounding it: only about 14% of Phase I applications are funded, and only about 40% of Phase I awardees advance to Phase II. Compounding those rates, roughly 5–6% of applicants reach Phase II. A financial model that treats a Phase II award as a planning assumption rather than a low-probability option is not a model. The grant path also imposes a real hidden cost — proposal writing, reporting and audit compliance consume meaningful founder and technical-staff time that is not on any budget line.

The second break point is the pilot-to-production gap. In the typical column, revenue at month 48 is $1.4M at 31% gross margin — real, but at pilot scale, where unit costs are far above the eventual production target. The company must now raise a much larger round to build production capacity, on the strength of three customers and thin margins. The "valley of death" between technical validation and commercial scale is a capital-markets problem, not a technical one.

Honest framing. The non-dilutive capital in this model is 18% of total — helpful, meaningfully so at the earliest stage when equity is most expensive, but grants do not finance a company. They de-risk specific technical questions and provide third-party validation. A deep-tech founder planning to reach commercial scale on grants alone is planning to spend ten years at sub-scale.


Cross-model comparison

Micro-SaaS VC B2B SaaS Consumer app AI startup Hardware Marketplace Regulated Deep tech
Capital to first $1M revenue ~$20K ~$2.5M ~$3.5M ~$2.0M ~$900K ~$2.8M ~$4.5M ~$7M+
Gross margin at scale 85% 76% 72% 44–58% 43% 54–61%* 64–71% 31–48%
Months to positive contribution 1 14–22 Never (typical) 15–18 1 (unit) / 30 (cash) 9–14 26 40+
Dominant failure mode Churn ceiling Burn multiple Paid UA arithmetic Negative GM at scale Working capital Liquidity per market Pre-revenue cost layer Funding gap
Most diagnostic metric Steady-state ceiling Burn multiple Revenue per install ÷ CPI Inference efficiency ratio Cash conversion cycle Liquidity by city CAC payback Months of runway to next milestone

* Marketplace margin shown as % of net revenue, not GMV.


PART C — WHICH METRICS MATTER AT WHICH STAGE

Most metric dysfunction comes from measuring the right thing at the wrong time. A pre-product-market-fit company tracking CAC payback is measuring the efficiency of a machine it has not yet built; a $30M ARR company that still cannot produce a cohort retention curve has a reporting problem that will cost it at diligence.

Pre-product-market-fit (0 → ~$1M ARR). Two metrics, and everything else is a distraction. Retention curve shape — does a cohort's usage flatten, or decay to zero? — and evidence of pull: are people paying without being sold to, and are they coming back unprompted? Growth rate is noise at this scale; a company going from 4 customers to 12 has grown 200% and learned nothing generalizable. CAC is unmeasurable because founder time is the acquisition channel and it is not in the P&L. Runway is the only financial metric that matters, because the only real question is whether you have enough months to find the thing.

Early product-market fit ($1M → $5M ARR). Now the machine exists and efficiency becomes measurable. Track gross revenue retention (not net — you need to know if the base leaks before you celebrate expansion), CAC payback on a fully-loaded basis, gross margin with COGS properly classified, and growth rate against the relevant benchmark — roughly 40% median at $1–5M for traditional B2B SaaS, materially higher for AI-native. Burn multiple becomes meaningful here for the first time. This is also the stage to fix metric definitions permanently, because restating history later is painful and looks bad.

Scaling ($5M → $25M ARR). NRR and GRR together, segmented. Burn multiple as the primary capital-efficiency metric. Magic number as the sales-specific version. Sales efficiency by segment and channel, because blended numbers now hide more than they reveal. ARR per employee as a discipline check. Cohort economics by acquisition channel, because the channel mix that got you to $5M will not get you to $25M at the same CAC.

Growth stage ($25M+). Rule of 40 becomes meaningful (and not before). Free cash flow margin. NRR by cohort vintage, which reveals whether recent customers are as good as older ones — the earliest warning that the market is saturating or the product is commoditizing. Segment-level contribution margin, because at this scale companies almost always contain one profitable business subsidizing one unprofitable one, and nobody knows which is which until they look.

Across all stages. Runway. Always runway.

Model-specific overlays. Marketplaces: liquidity per city-category before GMV, at every stage. Consumer: retention curve and revenue per install before anything else. AI-native: gross margin with inference disclosed, and committed-vs-consumption revenue split, from day one — because these are the metrics that determine whether the business is a software business. Hardware: cash conversion cycle, which no software metric captures. Deep tech: milestone progress and months to next de-risking event; conventional SaaS ratios are not merely unhelpful but actively misleading.


PART D — THE MOST COMMON METRIC-DRIVEN SELF-DECEPTIONS

These are listed in rough order of how much damage they do.

1. Annualizing the best month. Taking December's revenue, multiplying by twelve, and calling it ARR. Everything downstream — valuation, hiring plan, runway — inherits the error. The fix: trailing twelve months, always, alongside any run-rate figure.

2. Confusing gross flow with revenue. Reporting GMV as revenue in a marketplace, or the full customer payment as revenue when you are effectively reselling someone else's capacity. The ASC 606 principal/agent distinction is the formal version; the intuitive version is what would be left if the party you pass money to raised their prices 20%?

3. Reporting NRR without GRR. NRR of 118% is a very different company depending on whether GRR is 95% or 78%. The first is a product people keep and expand. The second is a product most people abandon while a few whales grow. The gap between NRR and GRR is the most informative number in startup reporting and the most frequently omitted.

4. LTV:CAC theater. Dividing by a churn rate measured over six months to produce a number describing five years, then comparing it to a CAC that excludes salaries. A 3x ratio built this way is commonly a 0.8x ratio built honestly. Use cohort payback instead.

5. Vanity denominators. Signups, downloads, registered users, "communities," pipeline, waitlist size. These correlate with effort, not value. The test: if this number doubled and nothing else changed, would the business be worth more? For most top-of-funnel counts the honest answer is no.

6. Calculating runway by division. Cash ÷ last month's burn is wrong whenever the plan involves hiring, whenever collections are lumpy from annual prepayments, or whenever there was an unusual expense. Model it forward under three scenarios and quote the middle one.

7. Survivorship-biased self-comparison. Benchmarking against a panel of funded survivors and concluding you are below average, then making aggressive decisions to close a gap against a cohort you were never in. If you are at the 25th percentile of a published SaaS benchmark, you are probably near the median of all companies that attempted the same thing. That is information, not an emergency.

8. Optimizing the measured thing instead of the real thing. Activation rates improved by redefining activation. DAU/MAU improved by notification spam. Conversion improved by lowering intent. Churn improved by excluding early churners. Every one of these makes a dashboard better and a company worse. The defensive practice: lock metric definitions, restate history when you change them, and pair every efficiency metric with a quality metric.

9. Mistaking a subsidy for a business. Negative gross margins, deep promotional discounting, supply-side incentives that never decline. These can be legitimate investments in a specific, dated mechanism — a cost curve, a density effect, a pricing migration. They are not legitimate as a permanent condition. The question is not "are you subsidizing?" but "what is the named, dated mechanism by which you stop, and what evidence do you have that it works?"

10. Treating benchmark medians as targets. The median company in any benchmark panel is not succeeding; it is surviving. Median NRR of 102% is not a goal, it is a description of the middle of a distribution that includes many companies quietly failing. Benchmarks tell you where you sit relative to a biased sample. They do not tell you what is good enough for your particular business, which depends on your capital, your market, and what you are trying to build.

11. Averaging across segments that should never be averaged. A $400 self-serve CAC and a $60,000 enterprise CAC blend into a meaningless $3,000. A profitable core product and an unprofitable new bet blend into a mediocre whole. Nearly every important operating insight lives in a segment, and blending is how it gets hidden — often from the founders themselves.

12. Believing the model. A spreadsheet with 40 assumptions produces an answer with 40 sources of error. The purpose of a financial model is not to predict the future; it is to make your assumptions explicit enough that you can find out which ones are wrong, and to identify the two or three variables the outcome is actually sensitive to. In every model in Part B, the ending value was dominated by one or two inputs — growth rate, churn, inference cost per task, cash conversion cycle. Find yours, instrument it, and stop polishing the rest.


Sources

Benchmark reports and data panels:

Metric definitions, frameworks and analysis:

Cost and pricing inputs used in the illustrative models:

A note on source quality. The SaaS Capital, Benchmarkit/Aleph, ChartMogul, RevenueCat, Carta and SPI/Deltek panels are primary data collections with disclosed sample sizes, and are weighted most heavily above. a16z, Bessemer and Craft Ventures pieces are investor analyses reflecting their portfolios and theses. Several cost inputs (money transmission, HIPAA, hardware tooling, burn-multiple-by-stage) come from practitioner guides rather than surveys and are cited as ranges, not point estimates. Where two sources disagree — as SaaS Capital and Benchmarkit do on ARR per employee, by 37% — both figures are shown rather than reconciled, because the disagreement is itself the finding.

Put the research to work.

All practical guides ↗

Find evidence

Design a paid pilot that can end cleanly

A credible pilot names the buyer, live scope, acceptance evidence, conversion decision, and exit work before implementation begins.

3 min · Practical guide

Find evidence

Run a concierge test without hiding the labor

Manual delivery can expose the real workflow, but only if the customer knows what is manual and the founder counts every minute required to deliver it.

4 min · Practical guide