# Unit Economics of Inference **Track:** Discovery & Validation at AI Speed — AI for Entrepreneurship — complete (29) **Framework / surface:** venture strategy **Level:** Intermediate **Prerequisites:** Pricing AI Outcomes **In one line:** Gross margin when COGS is tokens — and why the falling curve forgives sins it shouldn’t excuse. ## Theory, aesthetics & inspiration Martin Casado and Matt Bornstein's "The New Business of AI" (a16z, 2020) warned early that AI companies were shipping software economics with services-shaped costs: gross margins dragged into the 50–60 percent range, well below the 60–80-plus benchmark of comparable SaaS, by inference bills, human review, and per-customer variance. The response is an engineering discipline unit economics has never had to include before — caching repeated work, routing easy queries to small models and hard ones to frontier models, distilling expensive capability into cheap specialized weights, and batching where latency allows. The falling cost curve forgives much of this over time, and that is precisely its danger: a business that is only viable because tokens got cheaper has no moat against competitors enjoying the same discount. Model the margin at today's prices, treat the curve as upside, and know your cost per successful outcome — not per API call — because failures and retries are part of the true unit. **Founder question:** What does one successful outcome cost you, retries and review included?