SB StartupBasket
All ideas
77 /100 GO Low complexity

SpendVerdict — AI coding ROI referee for 20–150 dev teams

Tells an engineering leader which developers' AI token spend turned into shipped code, and which just burned budget.

— views
Evaluation Scores
77/100

GO

Overall Score

17
Problem
13
Demand
13
Build
12
Distrib.
11
Revenue
8
Time
3
Defense

SpendVerdict

1. One-liner

Tells an engineering leader which developers’ AI token spend turned into shipped code, and which just burned budget.

2. Trend signal — why now?

The token bill came due in 2026, and it came due loudly.

TechCrunch’s June 2026 investigation into runaway AI costs collected the quotes that define this moment. J.R. Storment, Executive Director of the FinOps Foundation: “In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our entire 2026 token budget and it’s only April.’” And on the shift in posture: “The whole conversation shifted from tokenmaxxing and ‘go fast’ to ‘we need guardrails, how do we control this?’”

Chris Reed, Senior Director of IT Finance at Priceline, on vendor dynamics: “It’s like the crack-cocaine epidemic. They let you try it to get you hooked on it, and now you’re kind of beholden to it.” His Cursor renewal came back “4-5x more expensive” and he reports “seeing discrepancies” between vendor-reported usage and internal data.

But the quote that actually describes the product is from Vitaly Gordon, CEO of Faros AI:

“One of my engineers spent $40,000 on tokens last month, and I genuinely don’t know whether I should stop him or should I go and tell everyone else to be like him.”

That is not a cost-control problem. Cost control is solved — LiteLLM will cap a budget, ccusage will print a number. That is a judgment problem, and nobody at this company’s size is selling the judgment.

The scale is now unmissable. Uber gave roughly 5,000 engineers Claude Code access in December 2025; usage nearly doubled by February 2026, 84% of developers were agentic-coding users by March, and the company burned its entire 2026 AI budget by April. Microsoft revoked its developers’ Claude Code licenses months after enabling them. Peter Steinberger ran ~100 Codex instances on one project and generated 603 billion tokens, 7.6 million requests, and over $1.3M in API spend in 30 days.

Meanwhile the per-developer number stopped being rounding error. DX’s 2026 pricing analysis: “The total cost per engineer, seat license plus token spend, is typically $200–$600/month for teams mixing inline and agentic tools” — which is $400,000–$600,000 annually for a 100-developer organization, before background API costs. The Pragmatic Engineer’s April 2026 survey of 900+ engineers found ~30% of engineers hit monthly usage limits and ~15% cite tool cost as a serious ongoing concern. One CTO in that survey: “we’re not sweating the costs because we’re trying to evolve best practices. But that has resulted in some devs really blowing through budget — so we may start instituting caps on spending.”

And the value side is genuinely ambiguous, which is what makes the judgment hard rather than obvious. Faros’s own 2026 AI Engineering Report found that under high AI adoption, bugs per developer are up 54%, the incident-to-PR ratio has more than tripled, median PR review time is up 441%, and code churn is up 861%. DX’s research across 400+ companies found median PR throughput improvement of just 7.76% (mean 13.1%, most organizations 5–15%). So spend went up 10× and output went up 8%. Somebody has to adjudicate that per-person, and right now nobody can.

The unlock that makes this buildable this quarter: Anthropic shipped the Claude Code Analytics API, exposing per-user sessions, lines of code, commits, pull requests, tool usage, and token/cost data by model via /v1/organizations/usage_report/claude_code, with data available for dates on or after 1 January 2026. The Admin API exposes usage_report and cost_report programmatically. GitHub, GitLab and Linear APIs have exposed the output side for a decade. The join is now a weekend of plumbing, not a data-engineering project.

Provenance:
  - Signal 1 (Demand): Named engineering and finance leaders on record that per-developer AI spend is unexplainable — Faros CEO Vitaly Gordon: "One of my engineers spent $40,000 on tokens last month, and I genuinely don't know whether I should stop him"; FinOps Foundation's J.R. Storment: companies "3x over our entire 2026 token budget and it's only April"; Priceline's Chris Reed on a Cursor renewal "4-5x more expensive". Uber burned its full 2026 AI budget by April. — https://techcrunch.com/2026/06/05/the-token-bill-comes-due-inside-the-industry-scramble-to-manage-ais-runaway-costs/ — 2026-06-05
  - Signal 2 (Feasibility): Anthropic's Claude Code Analytics API now returns per-user token/cost plus commits and PRs (data from 2026-01-01); Admin API exposes usage_report/cost_report; a review of the five leading Claude Code monitoring tools concludes none of them connect spend to shipped outcomes, and only the two free CLI tools serve teams under 50 devs. — https://platform.claude.com/docs/en/manage-claude/claude-code-analytics-api and https://www.toriihq.com/articles/five-claude-code-usage-dashboards-and-monitoring-tools — 2026-08-27
  - Signal 3 (Economic): All-in cost is $200–$600/engineer/month, i.e. $400K–$600K/year at 100 developers; the ROI-measurement category is real but priced for enterprise — Faros targets "500+ engineers, internal data team, $200k+ budget", Atlassian acquired DX for $1B in late 2025, LinearB runs $29–59/contributor. — https://getdx.com/blog/ai-coding-assistant-pricing/ and https://jellyfish.co/blog/faros-ai-alternatives/ — 2026-08-27
  Category: Tech-unlock

3. The opportunity

There are exactly two shelves in this market and the middle is bare.

The free shelf. ccusage parses local JSONL logs and prints a per-session cost report. Claude Code Usage Monitor forecasts when your billing window exhausts. LiteLLM enforces per-developer budgets at a gateway hop. These are good tools. They are also, respectively: single-machine only with “no team rollup or shared central dashboard”, “forecasts and live views only, no historical export”, and a piece of infrastructure you now have to run in front of every model call. All three answer “how much?” None answer “was it worth it?”

The enterprise shelf. Faros AI, Jellyfish, LinearB, and Atlassian-owned DX. These genuinely do connect AI spend to engineering outcomes — Faros’s own research on bug rates and PR throughput is the best public work on the question. But Faros is, per an independent 2026 review, “the right product for a narrow shape of customer: 500+ engineers, internal data team, custom-metric culture, $200k+ budget.” DX rolls 14 factors into a Developer Experience Index tied to business outcomes, which is a thing you buy when you have a director of engineering ops to operate it.

The company with 35 engineers spending $14,000 a month on Claude Code, Cursor and Codex is on neither shelf. Their VP Engineering has the exact problem Gordon articulated — one dev at $3K/month, another at $180, and no defensible way to say whether that’s brilliant leverage or an expensive habit. They will not deploy a gateway. They will not buy a $200K engineering intelligence platform to govern a $170K tool spend. They currently open the Anthropic Console, export a CSV, eyeball it next to a GitHub contributions graph, and guess.

The 10× is not another dashboard. It’s a verdict: a weekly, per-developer, defensible answer to “should I stop him, or should I tell everyone else to be like him?” — with the PRs, the review time, the revert rate, and the incident count sitting next to the dollar figure that bought them.

The reason this is fresh rather than obvious: until January 2026 the per-user cost data did not exist in an API. Vendors reported org-level aggregates. You could not attribute a dollar to a developer without running your own gateway, which is precisely why LiteLLM exists and precisely why small teams never did it. Anthropic opened that door eight months ago and the incumbents who walked through it were already pointed at the enterprise.

4. Target market

  • Primary customer: VP Engineering, Head of Platform, or a technical co-founder/CTO at a 20–150 developer software company — Series A through C, or a profitable bootstrapped product company. Global, English-speaking first: US, UK, EU, Canada, Australia, India. The buyer is the person who signs the tool renewal and gets asked by the CFO what it bought.
  • Why they buy: Their AI coding line item went from a rounding error to a top-five engineering expense in about nine months, and renewal season is here. They need to walk into a budget conversation with per-person evidence, and they need to know whether to cut seats, raise caps, or coach. The Pragmatic Engineer CTO quote is the customer speaking: “that has resulted in some devs really blowing through budget — so we may start instituting caps on spending.” Caps without evidence is how you cap your best engineer.
  • Rough TAM reasoning: Not a top-down TAM exercise — a bottom-up channel count. There are on the order of tens of thousands of software companies globally in the 20–150 engineer band with meaningful AI tooling adoption. I need ~700 of them at $250/month to hit $2M ARR. Anthropic’s Team plan alone starts at a 5-seat minimum at $100/seat/month, so the population paying enough to care is large and growing. This is a comfortably-under-$5M-ARR business and that’s the point.
  • Why now for them: Three things converged inside twelve months. Per-user cost data became API-accessible (Jan 2026). Spend crossed the threshold where finance asks questions ($200–$600/dev/month). And the first renewal cycle at the new prices landed — Priceline’s 4–5× renewal is the canary, and every VP Eng reading TechCrunch in June saw it.

5. Product sketch (MVP)

  • Connect in ten minutes: OAuth to GitHub (or GitLab), paste an Anthropic Admin API key, optionally connect Cursor/Copilot admin exports. No gateway, no proxy, no agent installed on developer machines.
  • The weekly verdict email. Every Monday, one page per team: each developer’s AI spend for the week sitting directly beside their merged PRs, lines changed net of reverts, PR review latency, and any incidents linked to their commits. Sorted by dollars-per-merged-PR.
  • Three buckets, named plainly. Each developer lands in Leverage (high spend, high shipped output), Burn (high spend, low shipped output), or Underused (low spend, and a seat you may be able to reclaim). The bucket comes with the two numbers that put them there.
  • Churn and revert flagging. Spend that produced code which was reverted or rewritten within 14 days gets subtracted from the “shipped” side, because Faros’s own data says churn is up 861% and counting it as output is how you lie to yourself.
  • Renewal pack. A one-click PDF for the CFO conversation: total spend by tool, cost per merged PR trended over six months, seats with zero usage in 30 days, and a recommended keep/cut/renegotiate line per tool.
  • Budget tripwires. Alert when any individual’s month-to-date spend crosses a threshold and their shipped output is in the bottom quartile — the combination, not the dollar figure alone. Nobody needs another spend alert.
  • Private by default. Individual-level views restricted to the leader who connected the account; team-level rollups shareable. This is a product that can very easily become a surveillance tool, and the ethics of that are addressed in section 13 rather than hand-waved.

6. AI angle — what’s load-bearing

Two places, and the product does not exist without either.

Attribution of output quality. Counting merged PRs is arithmetic. Deciding whether a PR represents real shipped value requires reading the diff: is this a genuine feature, a dependency bump, a formatting sweep, or 900 lines of AI-generated boilerplate that a reviewer will rewrite next sprint? An LLM classifying diffs into substantive-vs-mechanical, and detecting semantic reverts (the same logic rewritten by someone else two weeks later, which git does not record as a revert), is what separates this from a spreadsheet pivot. That classification is the entire value of the verdict.

The narrative. The deliverable is a paragraph a VP Eng can forward to a CFO: “Priya’s $3,100 this month produced 14 merged PRs across the billing migration with a 4% revert rate — that’s $221 per shipped PR against a team median of $780. Recommend raising her cap. Marcus’s $2,800 produced 3 PRs, two of which were reverted; recommend a conversation before renewal.” Generating that from the joined data, in language that survives contact with finance, is an LLM job.

Strip the AI out and you have a CSV join — which is exactly what these teams already do badly by hand, and exactly why they don’t do it.

7. Localization angle (if any)

N/A — this is a global play. The buyer persona (VP Eng at a venture-funded or bootstrapped software company), the tools (Claude Code, Cursor, Copilot), the billing currency (USD), and the language of the artifacts (English code review) are the same in Bangalore, Berlin and Austin. There is a mild pricing consideration — a $250/month tier is easier in the US than for an Indian services firm — but the primary market pays in dollars for the underlying tools already, so a single global price is correct. Forcing a localization angle here would be a mistake.

8. Business model — path to $1M–$5M ARR

  • Pricing: Flat per-workspace tiers, not per-seat — per-seat pricing on a tool that tells you to cut seats is a self-defeating incentive, and buyers see it immediately.
    • Starter — $149/month, up to 30 developers
    • Team — $299/month, up to 80 developers
    • Scale — $599/month, up to 200 developers, plus SSO and multi-tool consolidation
  • ACV: ~$3,600 blended (most customers land on Team). Against a customer whose AI tool spend is $80K–$400K/year, this is 1–4% of the budget it governs — the same ratio cloud FinOps tools charge, which is the anchor buyers already accept.
  • Rough math to $1M ARR: 280 customers × $299/mo × 12 = $1.0M. At a 20–150 dev company, that’s a single-decision-maker sale with no procurement committee.
  • Rough math to $5M ARR: ~1,150 customers at blended $3,600, or ~700 customers plus a move upmarket into the 200–500 dev band where a $1,200–2,000/month tier is defensible and where you’re now competing with Jellyfish on price from below. Realistically $5M requires both: broadening tool coverage (Copilot, Cursor, Gemini, Codex, Windsurf) so the product is the consolidation layer rather than a Claude-only add-on, and adding the 200–500 band.
  • Expansion path: Seat-band upgrades as teams grow; per-tool connectors as the customer’s stack sprawls; and the natural upsell is budget enforcement — once you’re trusted to judge the spend, selling the mechanism that acts on the judgment (per-developer caps applied through the vendor admin APIs) is a second SKU at roughly the same price again.

9. Go-to-market wedge — first 100 customers

  • The free teardown, run as a content engine. Build a free single-command tool that reads a team’s Anthropic Admin API export plus their GitHub org and prints the three-bucket verdict for one month, locally. Publish it. ccusage proved this exact audience adopts free CLI cost tools enthusiastically; the free tool is the top of funnel and the hosted weekly verdict is the paid conversion. Target: 3,000 CLI installs in 90 days, 4% converting to a paid workspace = 120 customers.
  • Cold outreach to the people who publicly complained. The TechCrunch piece, the Pragmatic Engineer survey, the FinOps Foundation Slack, and the r/ExperiencedDevs and Hacker News threads on AI tool costs are full of named engineering leaders describing this exact problem in June–August 2026. Build a list of 400 VPs Eng / CTOs at 20–150-dev companies who have posted about AI spend in the last six months. Send each a pre-computed teardown of publicly visible data (their org’s public GitHub PR velocity) plus a two-minute Loom on what the private half would add. Expect 25–30% reply on a personalized artifact like that, 5% close = ~20 customers.
  • Ride the renewal calendar. Cursor, Copilot and Claude Code renewals cluster at quarter and year end. Run a “Renewal Pack” campaign in the six weeks before each quarter close — the pitch is literally “walk into the renewal with per-person evidence.” Partner with the two or three independent FinOps consultancies now advertising AI cost practices; they need a tool to deliver, they have the buyer relationship, and a 20% referral share costs nothing until it works.
  • Publish the benchmark. After ~150 customers, publish an anonymized quarterly “cost per merged PR” benchmark by team size and tool. Faros and DX built their inbound on exactly this play and there is no small-team equivalent. This is the compounding channel — it makes the company the reference point for a number every VP Eng now needs.
  • Hacker News / Lobsters launch of the free tool. This audience is the HN front page audience, and “we measured whether your $40K/month of AI tokens shipped anything” is a title that lands there. One good HN day is 500–1,500 installs into the top of the funnel above.

10. Build complexity — justification

Low. Everything load-bearing is an off-the-shelf API: Anthropic’s Claude Code Analytics and Admin APIs return per-user cost and activity as JSON; GitHub and GitLab expose PRs, reviews, reverts and commit metadata; the LLM diff classification is a standard structured-output call. There is no gateway to run, no agent to install on developer machines, no data pipeline beyond a nightly pull into Postgres. The genuinely custom work is the classification prompt set and the revert/churn detection heuristics, plus the design of the weekly email — which is product work, not infrastructure work. A competent solo builder ships a credible v1 in 6–8 weeks; two people get to a polished multi-tool version in twelve. The dominant risk is not build difficulty, it’s connector maintenance as each vendor changes its admin API.

11. Gating checklist

GatePass?Note
Legal in target market✅Reads data the customer already owns, via official vendor APIs, with the customer’s own credentials. Employee-monitoring rules (esp. EU works councils, GDPR) apply and shape the product — see risk flags — but do not block it.
Ethical — no harm / dark patterns✅With a deliberate design constraint: individual-level data is visible only to the connecting leader, the product recommends coaching before cutting, and it explicitly flags “underused seat” rather than “unproductive person.” A version of this that ranks engineers publicly would fail this gate; that version is not what gets built.
Market exists (evidence above)✅Named executives on record; $400K–$600K/year spend at 100 devs; a funded enterprise category already selling the same insight upmarket.
1–5 person team can build this✅Solo builder, 6–8 weeks to v1.
Launchable with <$50K / ₹40L✅Realistically under $10K: hosting, LLM inference on diff classification, and a domain.

All five pass.

12. Feasibility score

AxisWeightScoreNotes
Problem intensity2017/20Felt weekly, costs real money now, and has a hard forcing function in renewal season. Not quite 18+ because a team can limp along with caps and vibes — it’s expensive ignorance, not a stoppage.
Demand evidence1513/15Multiple independent named-source signals: TechCrunch executives, Pragmatic Engineer’s 900+ respondent survey, a funded enterprise category (DX at $1B, Faros, Jellyfish, LinearB) proving people pay for exactly this insight. Docked 2 because no one has yet proven the small-team segment converts at $299.
Build feasibility1513/15Off-the-shelf APIs end to end, 6–8 weeks solo. Docked for multi-vendor connector sprawl.
Distribution clarity1512/15Free-CLI-to-paid is a proven motion in this exact audience (ccusage), the complainer list is real and nameable, and the renewal calendar is a genuine timing hook. Docked because free-CLI conversion rates are notoriously soft and 4% is an assumption.
Revenue mechanics1511/15Pricing is anchored to an accepted FinOps ratio and the ACV is honest. Docked because flat workspace pricing caps ACV, and $5M requires both upmarket movement and multi-tool coverage — two things that must go right, not one.
Time to first revenue108/10A single decision-maker, a credit card, and a problem they already know they have. Realistically 6–10 weeks from launch, with the free tool shortening it.
Defensibility103/10This is the weak axis and I won’t dress it up. The join is copyable in a quarter. Faros or DX could ship a self-serve tier down-market at any time, and LinearB is already priced close. The only real moats are the benchmark dataset (which compounds slowly) and being the default name before the incumbents notice the segment.
Total10077/100

13. Qualitative modifiers

Founder-fit tags

technical-heavy · content-heavy

This wants a builder who can ship API integrations fast and who can write — the free tool, the HN launch, and the quarterly benchmark are all content-led motions. It does not need enterprise sales chops, which is exactly why the segment is available.

Key assumptions to validate (3–5)

  1. Assumption: A VP Eng at a 20–150-dev company will pay $299/month for a verdict on a spend they already see in a console for free. How to test: Pre-sell. Take 40 named leaders from the complainer list, show a manually-produced teardown of their own data, and ask for a card on a founding-customer rate before the product is built. Need ≥8 to convert.
  2. Assumption: LLM diff classification can separate substantive from mechanical output accurately enough that engineers don’t dismiss the verdict as noise. How to test: Run the classifier over 500 merged PRs from 5 friendly repos, have the actual authors grade each classification, and require ≥85% agreement. Below that, the verdict is unusable and the product is dead.
  3. Assumption: Per-developer spend variance inside a single team is wide enough to be actionable. How to test: Pull the Anthropic Analytics API for 10 design-partner orgs. If the top spender isn’t at least 5× the median, there’s no decision to make and the product has no job. Gordon’s $40K anecdote suggests wide variance, but one anecdote isn’t a distribution.
  4. Assumption: The free CLI converts to paid at ≥3%. How to test: Ship the CLI first, standalone, and measure the rate at which installs request a hosted team rollup before the hosted product exists.

Risk flags

  1. Incumbent squeeze (highest). LinearB at $29–59/contributor is already within reach of this buyer, and Atlassian owns DX with Jira/Bitbucket distribution. If either ships a self-serve small-team AI-ROI tier, the differentiation collapses to price and speed. This is the reason defensibility scored 3.
  2. Platform dependency. The entire feasibility case rests on Anthropic’s Analytics API, plus equivalent access from Cursor, GitHub and others. Anthropic could restrict it to Enterprise plans, change the schema, or — worst case — ship this natively in the Console, which is a very plausible thing for them to do. Multi-tool coverage is the only hedge and it’s a partial one.
  3. Ethical and legal drift into surveillance. Per-developer productivity scoring is one product decision away from a tool that gets a founder pilloried on Hacker News, and in the EU it can trigger works-council and GDPR obligations around automated evaluation of employees. The mitigation is a real product constraint (leader-only individual views, coaching framing, no public leaderboards), and it costs some virality.
  4. Metric legitimacy. Engineers rightly distrust output metrics. If the verdict is perceived as counting lines of code with extra steps, adoption dies inside the team even if the VP bought it. The revert/churn subtraction exists specifically to answer this, and it has to be visibly correct.
  5. Market timing risk in both directions. If AI tool pricing stabilizes and per-dev cost drops toward $50/month, the problem shrinks below the pain threshold. If it keeps climbing, enterprises consolidate onto platforms that include governance and the standalone tool gets squeezed from above.

14. Structured verdict

Score:                  77/100
Verdict:                GO
Confidence:             Medium
Best-fit builder:       Technical solo founder who writes well; ex-platform or
                        dev-productivity engineer with a network of VPs Eng
Time to revenue:        6–10 weeks from launch
Capital to launch:      <$10K (hosting + LLM inference + domain)
Top 3 assumptions to validate first:
  1. Willingness to pay $299/mo — pre-sell to 40 named leaders from the
     public-complainer list, require ≥8 cards before building
  2. Diff-classification accuracy — 500 PRs across 5 repos graded by their
     own authors, require ≥85% agreement
  3. Intra-team spend variance ≥5× median — pull Analytics API for 10 design
     partners; if spend is flat across developers there is no decision to sell
Kill criteria:
  - Abandon if <8 of 40 pre-sell conversations convert to a paid founding slot
  - Abandon if diff classification lands below 85% author agreement after two
    prompt iterations — an untrusted verdict is worse than no verdict
  - Abandon if Anthropic, LinearB or DX ships a self-serve per-developer
    AI-ROI view for teams under 150 devs before v1 launches
  - Abandon if median design-partner top-spender is under 3× team median

15. Next step — 1-week validation sprint

  • Day 1–2: Build the list. 400 engineering leaders at 20–150-dev companies who posted publicly about AI coding tool costs between March and August 2026 — TechCrunch comment threads, the Pragmatic Engineer survey respondents’ public writing, HN threads on token spend, FinOps Foundation Slack, r/ExperiencedDevs. Cut to the 40 with the clearest stated pain and a reachable email.
  • Day 3–4: Produce five real teardowns by hand. Pick five design-partner-friendly orgs, get read-only Anthropic Admin API + GitHub access under NDA, and manually build the three-bucket verdict in a spreadsheet. This simultaneously tests assumption 3 (is variance ≥5×?) and produces the artifact for the outreach. Send the personalized teardown offer to all 40.
  • Day 5: Count. Go if ≥10 of 40 reply asking for the teardown AND ≥8 of the design-partner orgs show a top spender at ≥5× team median AND ≥3 verbally commit to $299/month at launch. No-go if replies are enthusiastic but spend variance is flat — that means the anecdote that inspired this is an outlier and there is no per-person decision to sell.

The falsifiable core: if the spend distribution inside real teams turns out to be uniform, there is no verdict to render and this idea is dead regardless of how loudly people complain about the total bill. That’s the number to go get first.

Interested in a detailed proposal?

Get a deep-dive with market research, competitive analysis, and implementation roadmap.

Contact us

info@startupbasket.ai