THE COMPANY-BUILDING FIELD NOTEBOOKRESEARCH EDITION / SEPTEMBER 2026
Startup
Research.
Search
RESEARCH LIBRARY / The starting point

The research, in perspective

Read the central findings, definitions, and limitations before following a particular path.

Startups: How to Build, Fund, Grow, and Scale a Company

Research date: September 15, 2026.

This document summarizes a fifteen-chapter research library totaling roughly 279,000 words, compiled to support a website about startups and entrepreneurship. Every figure below is sourced in the chapter it comes from. Where chapters disagree with conventional wisdom, they say so and show the data.


1. What this library is, and how it was built

The brief asked for deep, current research on the complete process of turning an idea into a sustainable company, serving an audience that ranges from first-time founders to investors to students. Fifteen chapters were produced:

# Chapter Words Cited sources
01 Startup ecosystem and taxonomy of company types ~11,900 43
02 Sector-by-sector analysis (26 sectors) ~15,400 75
03 Idea validation and product building ~12,400 64
04 Tools and platforms (22 categories, 8 costed stacks) ~12,400 71
05 Founders, teams, and go-to-market ~15,500 56
06 Metrics and unit economics (8 financial models) ~17,800 39
07 Funding and financing ~19,300 36
08 Accelerators, investors, and support systems ~15,200 60
09 Case studies (36 companies) ~22,000 89
10 Pitfalls and failure modes (+ two frameworks) ~14,800 36
11 Legal, operations, and founder sustainability ~25,200 75
12 Website strategy and roadmap ~17,200 30
13 Endgame and ownership strategy ~19,200 47
14 Exit deal mechanics ~22,100 22
15 Post-mortem library (53 failures) ~22,500 157

Methodological commitments applied throughout. Every chapter labels claims by epistemic status — verified fact, estimate, projection, founder claim, or the author's own analysis. Medians are used wherever averages are distorted by outliers, and the chapter says which it is using. Survivorship bias is flagged explicitly and repeatedly, including the inverse problem that failure analysis suffers from hindsight bias. Conflicts of interest are named: venture-data providers sell subscriptions and benefit from a narrative of a large trackable market; accelerators publish their own alumni statistics; market-research vendors sell reports and a bigger number sells better; VC-published content promotes raising and bootstrapper content promotes not raising. Nothing in the library is framed as personalized legal, tax, or investment advice, and Chapter 11 marks the specific moments that genuinely require a professional.

One methodological note worth surfacing. Where volatile items were involved — non-compete enforceability, Section 174, QSBS, AI regulation, accelerator terms, tool pricing — the chapters re-verified current status against primary sources rather than relying on recalled knowledge. Several had moved materially in the past twelve months, and are flagged below.


2. The fifteen findings that matter most

2.1 The venture path is statistically a rounding error, and most startup writing obscures this

Against roughly 5.2 million annual US business applications — or about 1.7 million "high-propensity" applications, the Census Bureau's better proxy for real firm formation — roughly 10,000 companies are on pace to raise a first institutional venture round in 2026 — which would be a record year. That is 0.19% of applications, or 0.57% on the more generous denominator (author's calculation; in a non-record year the share would be lower). Fewer than one new American business in a hundred and fifty ever raises venture capital.

Meanwhile 36.2 million US small businesses represent 99.9% of firms, and 82.3% of them are nonemployer firms — one person working for themselves. The overwhelming majority of American businesses look nothing like the companies startup media covers.

Implication for the site: a publication that treats the venture path as the default is describing a tiny, unrepresentative slice while implying it is the norm. The single most differentiated editorial position available is to report base rates accurately.

2.2 "90% of startups fail" has no traceable source, and the real numbers are different

Chapter 10 traced the figure and found no underlying study. The actual administrative data is BLS Business Employment Dynamics. Chapter 10 reports the March 2015 establishment cohort at 79.6% surviving one year, 50.2% five years, and 34.7% ten years. Chapter 01 tracks the older March 2005 cohort through a full twenty years: 80.1%, 46.8%, 33.8%, and 19.5% at one, five, ten, and twenty years — lower at the five-year mark because that mark fell inside the Great Recession. The two chapters measure different cohorts rather than contradicting each other. Roughly half of new establishments survive five years.

Two caveats the chapter insists on. "Survival" means the establishment exists and has employment; it says nothing about whether the owner earns a good living, and a business can survive twenty years while returning less than the owner would have made employed. And the failure-reason literature is mostly self-reported: CB Insights' March 2026 edition analyses 431 shut-down VC-backed companies, 385 of them with identifiable causes, 70% citing "ran out of capital." Chapter 10 argues at length that "no market need" functions as a category masking at least six distinct underlying failures, and is what founders say when the real cause is harder to admit.

2.3 AI absorbed seven to nine of every ten venture dollars — and the structural distortion matters more than the headline

Global venture funding hit $510 billion in H1 2026, a record half-year against $440 billion for all of 2025. OpenAI and Anthropic alone took $217 billion of it — 43% of the global total.

AI's share depends on who is counting: Crunchbase says over 70% of global Q2 capital; PitchBook-NVCA says 86% of all US H1 2026 venture dollars ($355.9B across 3,258 deals); CB Insights records 263 mega-rounds taking 81% of capital with deal count at its lowest in over a decade. These are not in conflict — they use different definitions of "AI company" and different geographies.

The more actionable distortion is structural. Megadeals of $100M+ took 87.5% of US capital, while first-time fund formation is on pace for its lowest year since 2016. First-time managers are historically the buyers of non-consensus early rounds. Record headline totals coexist with non-AI seed and Series A being genuinely harder to raise than the numbers suggest.

2.4 Sector verdicts: seven promising, six overcrowded, six hype, and a capital-starved category most analyses omit

Chapter 2's synthesis, in brief. Genuinely promising: cybersecurity (healthy seed activity, working exits, and the specific opening of identity and permissions for non-human agents); vertical SaaS with payments attach; healthcare administration (roughly a quarter of $5.7 trillion in US health spending); grid and interconnection software (Futurum estimates an ~$80 billion Azure order backlog constrained by power availability — an analyst figure the verification pass could not trace to a Microsoft disclosure, and which Chapter 02 now labels accordingly); construction and industrial productivity; defense supply chain beneath the primes; and biotech countercyclically.

Overcrowded: the AI application layer generally — RevenueCat data shows AI consumer apps convert 52% better but churn 30% faster with 20% higher refunds, a novelty-purchase profile; AI coding tools; humanoid robotics; consumer apps and creator tools; gaming studios (record profits alongside ~14,666 forecast layoffs); horizontal AI infrastructure below the frontier.

Mainly hype, defined narrowly as the narrative running ahead of demonstrated economics: general-purpose humanoids; orbital data centers; "agentic" as a product category; consumer web3 (concluded, not emerging — stablecoin rails are the real business); pure outcome-based SaaS pricing; AI drug-discovery valuations (the science, not the valuations — preclinical speed is not the binding constraint, Phase II efficacy is).

Capital-starved — decent fundamentals, no scale-up money: climate tech Series C down 32% to an all-time low while Series D+ rose 78%; agtech deeptech received a 78% seed valuation premium and zero $200M+ rounds; edtech; and accessibility tech, which no major venture data provider tracks as a category at all.

2.5 Solo founding has stopped being a disadvantage on the metrics that are measurable

Carta data shows the solo-founder share of new startups rising from 23.7% to 36.3% between 2019 and H1 2025. Solo founders were 30% of 2024 startups but captured 14.7% of capital — a real financing penalty, not necessarily an outcome penalty. Chapter 5 pairs this with academic work (Howell and Hall, SMJ 2026) on "T-shaped" founder capability.

Chapter 5 also corrects a piece of folklore directly: the widely-repeated claim that 65% of startups fail from co-founder conflict is not supported; Carta's data indicates a 25–35% five-year co-founder breakup rate. Conflict is common and consequential; the specific statistic is fabricated.

On equity, Carta's median split is 51–49, and Wasserman's research found 73% of teams split within a month of starting — before knowing what anyone would contribute.

2.6 AI-assisted development is real but the productivity claims are not established

Chapter 3 assembled the evidence rather than the marketing. METR's randomized controlled trial found experienced developers took 19% longer using AI tools on familiar codebases while believing they were faster — with the study's own caveats about generalizability. GitClear's analysis of 623 million code changes found maintainability degradation. Veracode found roughly 45% of AI-generated code contained vulnerabilities. Both of the latter come from vendors with an interest in the result, which the chapter flags.

The chapter also documents real incidents — the Replit database deletion, the Wiz/Base44 vulnerability — and explicitly debunks the widely-repeated claim that the Tea app breach was a vibe-coding failure. The honest summary is that AI-assisted development changes the cost structure of building software substantially, and that the second-order effects on security, maintainability, and defensibility are unresolved.

2.7 AI gross-margin compression is genuinely contested, not settled

Chapter 6 found the two best available panels disagree. The 2026 Aleph × Benchmarkit panel (342 companies, CY2025 actuals) could not detect margin compression at the median. High Alpha's 800-company panel found roughly 10 points of year-over-year compression among early-stage firms.

The chapter's Model 4 shows the mechanism concretely: an agentic workload costing $0.90 per resolution on frontier models and $0.23 after routing and caching produces −168% gross margins for heavy users under seat-based pricing even after optimization. The compression is real at the workload level; whether it shows up in a company's blended median depends on pricing design and user mix.

The same panel showed median gross revenue retention falling four points to 84% while the magic number jumped from 0.94 to 1.37 — efficiency improved while retention weakened, which is an uncomfortable combination.

2.8 SEO as a customer acquisition channel has structurally degraded

SparkToro measured 68% zero-click searches; Pew found AI-summary result pages produce 8% click-through versus 15% without. Chapter 5 treats this as a change in kind rather than degree for informational content, and Chapter 12 builds the entire website strategy around it.

This finding is load-bearing for the site itself, and Section 4 below returns to it.

2.9 Several things founders "know" about funding changed in the last eighteen months

Chapter 7's most operationally important verifications: SBIR/STTR lapsed in October 2025 and was reauthorized in April 2026 through 2031, with new $30M Strategic Breakthrough Awards — a gap that stranded applicants. Pipe exited standalone revenue-based financing. SEC data shows the median Reg CF raise is $113,000 with a 0.25% eventual IPO rate, against a marketing narrative of crowdfunding as a serious capital source.

Chapter 8 verified accelerator terms rather than repeating stale figures: Y Combinator now runs four batches a year (W27 deadline November 2, 2026), terms unchanged at $125K for 7% plus $375K uncapped MFN. Techstars moved to $220K for 5% common in April 2025, after closing Seattle and DC and closing then later reviving Boulder. Where current terms could not be verified — Antler, Entrepreneur First, Alchemist, HF0, 500 Global — the chapter says so rather than restating old numbers.

Chapter 11 re-verified every volatile item. The FTC withdrew its non-compete appeals in September 2025. Section 174A restored domestic R&E expensing, with a retroactive election deadline of July 6, 2026. QSBS moved to tiered exclusions with $15M/$75M caps for stock acquired after July 4, 2025. Colorado's AI Act was repealed and replaced by SB 26-189, effective January 1, 2027. The EU AI Act's Annex III high-risk obligations were deferred to December 2, 2027, while Article 50 transparency duties hold at August 2, 2026. The FTC's click-to-cancel rule was vacated in July 2025, with a new ANPRM pending. COPPA's amended compliance date was April 22, 2026, and CPPA ADMT rules take effect January 1, 2027.

Enforcement is not theoretical. Chapter 10 documents the cases: Healthline was fined $1.55M and Tractor Supply $1.35M under California privacy law; Anthropic settled copyright litigation for $1.5 billion; Moffatt v. Air Canada established that a company is responsible for what its chatbot tells a customer.

2.11 The cost of starting has collapsed; the cost of being noticed has not

Chapter 4's eight costed stacks show a solo founder can reach production on a few tens of dollars a month, with generous startup credit programs available from AWS, Google, Microsoft, and Nvidia. Chapter 4 also verified two status changes founders still get wrong: Bench collapsed on December 27, 2024 ($2.8M cash against $65.4M liabilities) and was acquired by Employer.com three days later, and Stripe acquired Lemon Squeezy in July 2024, which now steers users toward Stripe Managed Payments.

The asymmetry this creates is the through-line of the whole library. Building is cheap and getting cheaper; distribution is expensive and getting more expensive. Chapter 3's defensibility section and Chapter 5's channel analysis both conclude that the durable advantages available to a new company are distribution, proprietary data, workflow depth, and regulatory position — not the software itself.

2.12 Case studies are the most misleading genre in startup literature

Chapter 9 opens with an extended methodological argument before presenting a single company: base rates from Census BFS, BLS, and Ghosh; Wald's survivorship problem; hindsight bias; and the fact that founder retrospectives are reconstructed narratives rather than records, frequently embellished in documented ways.

The 36 cases were deliberately outcome-mixed — including modest successes like TinyPilot's $598K exit and outright failures and frauds — rather than a unicorn parade. Status changes verified as of September 2026 include Nvidia's $12.93B acquisition of Hugging Face (announced September 3, 2026), Figma down roughly 80% from its August 2025 peak (at $24 as of February 2026), Bumble at a $417M market cap, 23andMe now a Wojcicki-owned nonprofit, and Ginkgo Bioworks at $20M quarterly revenue after a 1-for-40 reverse split.

The chapter closes with patterns that hold up across the set, and eight commonly-claimed patterns that do not.

2.13 Earn-outs pay roughly twenty cents on the dollar

Chapter 13 and Chapter 14 converge on the single most under-known number in exits. SRS Acquiom's framing of its own escrow-agent data is that "closer to one out of five dollars gets paid" across all deals carrying an earn-out, and roughly half of earn-out deals pay nothing at all.

The practical consequence: an offer of "$40M — $25M at close and $15M in earn-out" has an honest expected value near $28M, not $40M. Founders routinely accept a lower cash number in exchange for a larger earn-out and are systematically worse off for it. Chapter 14 covers why — the measurement-period problem, and the fact that post-close the buyer controls the very operations the targets are measured against.

2.14 The headline sale price is often unrelated to what the founder receives

Chapter 14 works two full proceeds waterfalls, arithmetic verified line by line. In the clean case — a bootstrapped $12M sale — a 60% owner nets about 38% of the headline after banker fees, working-capital adjustment, escrow, and tax.

In the venture case, a "$75M" sale leaves a 12% founder with roughly $28,000 from equity, with a further ~$735,000 arriving only through a discretionary compensation carve-out the buyer and preferred stockholders agreed to fund. The preference stack consumed almost everything. Chapter 13 works a parallel $100M sale that leaves founders about $1.4M each.

A related finding worth surfacing early: Chapter 13's analysis shows a 1x non-participating preference stops binding much sooner than founders assume. In its seed-stage worked example, the preferred converts above roughly a $23M sale price — after which the founders' full ownership percentage applies to the whole price.

2.15 What founders say killed their company differs from the evidence 45% of the time — and the gap widens with capital raised

Chapter 15's original contribution. Across 53 structured post-mortems, 44 carried an attributable stated cause. Of those, 20 (45%) diverge materially from what the record supports — and the divergence is directional: ten shifted blame outward, six reframed a failure as something other than failure, four softened it, and none was harsher on the founder than the evidence warranted.

The divergence correlates strongly with money raised: under $10M, the stated cause matched the evidence in 14 of 17 cases (82%); over $100M, in 4 of 18 (22%). The more capital involved, the less reliable the founder's account.

The chapter is candid that its sample is biased toward failures someone chose to write up — which skews venture-backed, and skews toward founders whose failure was instructive rather than humiliating. The failures nobody documents are absent by construction.


3. Cross-cutting theses

Four arguments recur across chapters and are, in combination, the library's intellectual position.

Base rates before narratives. Nearly every widely-repeated startup statistic is either unsourced, methodologically weak, or describing a different population than the reader assumes. The 90% failure rate, the 65% co-founder conflict figure, Startup Genome's 74% premature-scaling claim — all are examined and either corrected or heavily qualified. A reader who internalizes the actual base rates will make better decisions than one who has read ten times as much conventional startup content.

The financing decision is a choice about what kind of life you want, not a measure of ambition. Chapter 1's taxonomy and Chapter 7's closing comparison both arrive here. Venture capital is a specialized instrument suited to a specialized case: businesses that can plausibly return a fund. Fifteen other company models exist, most with better median outcomes and worse maximum outcomes. The relevant question is not "can I raise?" but "does the shape of this business match the shape of the instrument?"

Asymmetric cost curves have changed which advantages are durable. Production costs are collapsing across software, content, and increasingly design and research. Attention costs are rising. This inverts the traditional sequence in which a founder builds something good and then finds customers. Chapters 3, 5, and 12 independently conclude that distribution should be solved before or alongside the product, not after it.

The endgame is chosen years before anyone calls it an exit. Chapter 13's central argument, and the reason it sits alongside the funding chapter rather than after it. Entity type, who is on the cap table, and which instrument a founder took quietly foreclose most of the menu of destinations long before an exit is discussed. Three destinations close at the first priced round. A founder who has not decided where they are going will have the decision made for them by the structure they accepted early.

Honest uncertainty is more useful than false confidence. The library repeatedly reports that credible sources disagree — on AI margin compression, on AI's funding share, on whether accelerators add value net of selection effects — rather than manufacturing a clean answer. This is both intellectually correct and, per Chapter 12, the site's most defensible market position.


4. What the research implies for the website

Chapter 12 is deliberately unflattering about the opportunity, and its conclusions should be read before any building begins.

The category is among the worst-positioned that exists. Generic startup advice is exactly the content type that AI answers have absorbed. The space is also densely occupied — Crunchbase, PitchBook, TechCrunch, Y Combinator's own library, First Round Review, Lenny's Newsletter, Indie Hackers, SaaStr, Failory, Starter Story, CB Insights, Carta's data blog — much of it by organizations with data assets or capital a solo operator cannot match.

The recommended position is "the base-rate desk." Not a startup advice site: a place that reports what actually happens, with sources and with medians, and that corrects the numbers everyone else repeats. This is the one position the research library is uniquely equipped to occupy, because producing it required exactly the work already done.

SEO is treated as dead upside rather than strategy. The distribution plan does not assume search traffic.

Several requested features should be cut or scoped down. A funding-round tracker and a startup directory both require ongoing data pipelines that are expensive to maintain and are already dominated by well-capitalized incumbents — and Crunchbase's API access for this use case has been discontinued. The tool directory is reduced to eight curated stacks. Calculators rank first among proposed surfaces on the value-to-effort-to-defensibility assessment: runway, dilution, and pricing calculators are cheap to build, genuinely useful, linkable, and hard for an AI answer to replace.

Revenue expectations are modeled honestly. The most likely year-one outcome in Chapter 12's model is roughly $1,000. The chapter includes current affiliate program terms for five verified startup tools and niche B2B newsletter sponsorship benchmarks, and is explicit about what fraction of such sites earn meaningfully.

Chapter 12 also delivers the requested operational artifacts: a full information architecture with URL conventions and topic-cluster strategy, 60–80 suggested page titles, a publishable editorial standards document (including AI-assistance disclosure and affiliate disclosure per FTC rules), a data methodology covering which public datasets are usable and under what licensing (SEC EDGAR and Census are viable; Crunchbase is not), a scoped MVP, a week-by-week 90-day launch plan, and a 12-month roadmap with effort estimates.


5. Practical frameworks available for immediate use

Three artifacts in the library are directly publishable with light editing:

The "should we build this?" decision framework (Chapter 10) — a four-gate, six-dimension structure covering problem severity, market size and reachability, willingness to pay, founder fit, distribution advantage, capital requirements, regulatory exposure, defensibility, and timing, with worked examples of a passing, a failing, and a borderline idea.

The staged startup risk checklist (Chapter 10) — 40 diagnostic questions organized across pre-idea, pre-launch, post-launch, and scaling stages, each with an explanation of what a "no" implies.

The validation ladder (Chapter 3) — the distinction between interest, usage, willingness to pay, and product-market fit, with what each validation method actually proves, what it does not, and the specific self-deception each one invites.

The endgame decision framework (Chapter 13) — a seven-step sequence taking a founder from their business type, cap table, financial needs, and time horizon to the two or three endgames actually available to them, with what would have to change to open up others, worked through four different starting positions.

The pre-sale readiness audit (Chapter 14) — the diligence problems that kill deals late, which a founder should be fixing twelve to twenty-four months before any conversation with a buyer.

Chapter 6's eight financial models, Chapter 4's eight costed tool stacks, and Chapter 15's 53-entry structured post-mortem database are similarly ready to become interactive site features.


6. Known limitations

Stated plainly, because the library's own standards require it.

Venture data is revised upward for several quarters after a period closes, so H1 2026 figures cited here will likely be restated higher. Private valuations are negotiated rather than observed, and credible sources differ by meaningful margins. Benchmark panels overwhelmingly sample surviving, funded companies, which biases every retention, growth, and efficiency figure optimistically. Tool pricing and accelerator terms change frequently and several were verified within days of this writing. Regulatory items in flux — AI regulation especially — have effective dates that may shift again. Two chapters exhausted their web-search budgets and completed verification through direct fetches of primary sources, which is methodologically sound but means fewer alternative sources were canvassed for a minority of claims.

An independent verification pass audited the highest-risk claims across all chapters against primary sources. It confirmed the Y Combinator and Techstars terms, the SBIR lapse and reauthorization, the AI funding figures, the Nvidia/Hugging Face acquisition, all seven Chapter 11 regulatory dates, the Carta founder data, the METR and Veracode findings, and the Bench and Lemon Squeezy status changes. It found two factual errors — two of the four BLS survival figures in Chapter 01, repeated in this summary — plus ten lesser issues of attribution, labeling, and staleness.

A second pass audited chapters 13–15, with priority on their numerical content. Both of Chapter 14's proceeds waterfalls verified exact on every line, across roughly forty independent checks, as did Chapter 13's $100M waterfall and three of its four worked examples. That pass found five errors: an exit threshold in Chapter 13 overstated by about 85%, an internally inconsistent earn-out decomposition, and three factual slips in Chapter 15 including two tabulations that did not reconcile to the chapter's own entry count. All were corrected, and Chapter 15's synthesis was rebuilt from an auditable per-entry table — which moved its headline divergence finding from 54% to 45%.

All twenty-two issues across both passes have been corrected in the files as delivered. The full audits are in VERIFICATION-REPORT.md and VERIFICATION-REPORT-2.md, retained deliberately so the corrections are auditable rather than silent.

Nothing in this library is legal, tax, financial, or investment advice. Chapter 11 identifies the specific decisions that require a qualified professional, and that list should be treated as the minimum.


Compiled September 15, 2026. Chapter files 01–12 accompany this summary; each carries its own full source list.