SB StartupBasket
All ideas
74 /100 GO Medium complexity

LiftReferee — delivery ad referee for independent restaurants

Runs the holdout test delivery apps will never run, and tells you which sponsored orders you were getting anyway.

— views
Evaluation Scores
74/100

GO

Overall Score

16
Problem
13
Demand
11
Build
12
Distrib.
11
Revenue
8
Time
3
Defense

LiftReferee — delivery ad referee for independent restaurants

1. One-liner

Runs the holdout test delivery apps will never run, and tells you which sponsored orders you were getting anyway.

2. Trend signal — why now?

Three things happened in the last eighteen months and they point at the same hole.

Delivery advertising stopped being a rounding error. DoorDash, Uber Eats and Instacart now generate over $4 billion in annualised advertising revenue between them — DoorDash past $1B, Uber Advertising past $2B on a run rate, Instacart at $286M in Q1 2026 alone (Osmos, 2026). Uber management told the Q1 2026 earnings call that delivery ad penetration has crossed 2% of gross bookings. DoorDash’s Q4 2025 ad revenue grew 73% year-over-year to $780M, now ~12% of total revenue, up from 4% in 2022 (OPA!). That money comes out of restaurant P&Ls.

Ad spend has quietly overtaken commission. The same analysis documents a 25-location fast-casual chain whose DoorDash ad costs reached 18% of gross sales — nearly double their 10% commission rate. Typical operators now run 8–15% of delivery sales on ads on top of ~25% commission, putting total platform cost at 35–40% of delivery sales. Operators watch the commission line, which is contractual and visible, while the ad line — which they nominally control — grows past it unnoticed.

Nobody is measuring incrementality for the little guy. Platforms self-report ROAS on a 7-day click-attributed window; some competitors attribute on view-through (search summary of platform attribution norms). Industry commentary is explicit that these models “credit conversions to ad exposure without accounting for whether those purchases would have occurred regardless of advertising, which can lead to inflated performance metrics.” DoorDash itself now promotes Ghost Ads incrementality testing — in eight large-scale tests it “reduced measurement noise by over 90%” — but that’s a CPG-enterprise product, not something an independent taqueria gets. Meanwhile the tools that do run proper incrementality — Pacvue, Skai, Flywheel, CommerceIQ — are the named DoorDash Ads API partners, and Pacvue is openly aimed at brands spending $50,000+ per month; “if you’re spending less than $50K/month on retail media, Pacvue is likely overkill” (atom11).

The restaurant spending $3,000 a month sits below every incumbent’s floor and above the point where the waste stops being trivial.

Provenance:
  - Signal 1 (demand): DoorDash ad spend reached 18% of gross sales at a 25-location chain — nearly double its 10% commission — while operators track commission and miss ads; typical operators now at 8–15% of delivery sales, 35–40% all-in — https://opalink.com/blog/the-hidden-cost-of-doordash-advertising-how-restaurant-ad-spend-now-exceeds-delivery-commissions — observed 2026-08-31
  - Signal 2 (feasibility): DoorDash Ads Manager Report Builder exports daily-level campaign data as CSV including ROAS, orders, AOV, ad fee, CPA and a new-vs-existing customer split, self-serve in the Merchant Portal; Uber Eats Manager exports CSV reports the same way — no gated API required — https://help.doordash.com/en-us/advertising/article/ads-manager-report-builder and https://help.uber.com/en/merchants-and-restaurants/article/how-can-i-view-my-performance-data-in-uber-eats-manager — observed 2026-08-31
  - Signal 3 (economic): DoorDash + Uber Eats + Instacart past $4B annualised ad revenue (Uber past $2B run rate, >2% of delivery gross bookings); combined US delivery-app ad run rate likely to cross ~$5B annualised before end-2026 — https://www.osmos.ai/blog/food-delivery-retail-media-monetization-2026 — observed 2026-08-31
  - Signal 4 (incumbent floor): Pacvue and peers target brands spending $50,000+/month on retail media and are "likely overkill" below that — the sub-$10K/mo restaurant advertiser is explicitly unserved — https://www.atom11.co/blog/pacvue-pricing-guide — observed 2026-08-31
  Category: Underserved niche

3. The opportunity

Every party in this market is paid to tell the restaurant its ads worked.

The platform grades its own homework: DoorDash reports ROAS on a 7-day click window, meaning a regular who was going to order your pad thai anyway, who happened to tap a sponsored tile on the way, becomes attributed ad revenue. Higher reported ROAS means more budget next month. That’s not fraud — it’s the standard retail-media attribution model — but the platform has no incentive to build the counterfactual.

The ad tools sold to restaurants — Otter Marketing, Revly, Chowly — are optimisation layers. Otter claims “ROI increases of up to 10–20%” and automates campaigns. Revly is “an all-in-one platform integrating ads, offers and data insights.” Both are excellent at running ads. Neither page mentions incrementality, holdout groups, or controlled experimentation (Otter, Revly). Structurally they can’t: a product whose value proposition is better ad performance cannot credibly deliver the verdict “turn 60% of this off.” That’s the vendor conflict of interest, and it’s the gap.

The enterprise incrementality vendors who do this properly won’t come downmarket. Their floor is $50K/mo of media spend. A restaurant spending $3K/mo is not a customer they will ever chase.

So the independent operator is left with a number they can’t trust, produced by the party selling them the thing, with no cheap way to check it. Meanwhile the check is not hard. The platform hands over daily campaign data including a new-versus-existing customer split. The operator’s POS holds total channel sales. Turn ads off in a disciplined, pre-registered pattern — specific dayparts, specific weeks — hold everything else constant, and the difference between forecast and actual baseline sales is the incremental lift. This is a 1930s agricultural experiment design applied to a $5B ad market that never bothered.

The wedge is that LiftReferee sells the opposite outcome to everyone else in the category: the number that lets you spend less.

4. Target market

Primary customer: Owner-operator or marketing lead at an independent restaurant or small chain, 1–20 locations, US and UK, doing $500K–$5M annual revenue per location, with $900–$8,000/month of combined DoorDash + Uber Eats ad and promotion spend. Sweet spot: 2–8 locations, where the spend is big enough to hurt ($5K–$25K/mo) but the operator is still the person reading the reports at 11pm.

Why they buy, in the terms they use: the industry framing is that for most restaurant operators, ad spend on delivery platforms is a black box — “they put money in, see some orders come out, and hope the math works.” Restaurant marketing guidance for 2026 says the same thing about the wider budget: operators “struggle to track marketing effectiveness, guessing what worked when promotions run through channels outside their own systems,” and that marketing “can no longer be treated as a guessing game” (Total Food). The pressure is real: an Omaha operator paid $188,000 in one year of delivery app fees before dropping the apps entirely (Moneywise). Operators are not confused about whether the platforms are expensive. They are confused about which specific line they can cut without losing volume — and that uncertainty is exactly what keeps them paying.

There’s a second, sharper version of this fear. Commentary on the category notes DoorDash’s algorithm changes through 2025 systematically reduced organic visibility for non-advertising restaurants. So the operator’s real question isn’t “do ads work” — it’s “if I stop, does the algorithm bury me?” That is a genuinely empirical question and nobody is answering it for them.

Rough TAM reasoning: DoorDash reports its unified global commerce media platform carries more than 400,000 advertisers. Restaurants are a large share. If even 5% of those are independents or small chains spending enough to justify $199/mo, that’s 20,000 addressable accounts on one platform. Capture 1,500 of them and you’re at $3.6M ARR. This is a small-market play by design — too small for the enterprise incrementality vendors to defend, plenty for a bootstrapper.

Why now for them: ad spend crossed the line from marketing experiment to top-three cost line in the last two years, margins are compressed by 35–40% all-in platform take, and both major platforms now expose the daily campaign export that makes the analysis possible without begging for API access.

5. Product sketch (MVP)

  • Connect your reports, not your bank. Upload (or auto-pull via emailed scheduled report) the DoorDash Ads Manager Report Builder CSV and the Uber Eats Manager campaign CSV. Optional POS connection (Toast, Square) for total channel sales; without it, the platform’s own sales report is the baseline.
  • The Baseline Model. Twelve weeks of history in, and the product learns your normal: day-of-week shape, daypart curve, weather and holiday effects, promotional spikes. This is the counterfactual — what your delivery sales do when you leave them alone.
  • Pre-registered holdout scheduler. The product designs the test and tells you exactly what to switch off and when: “pause Sponsored Listings on DoorDash, Tuesdays and Wednesdays, 11am–2pm, for the next three weeks. Leave Uber Eats untouched as your control.” It writes the schedule; you flip the switches in the merchant portal.
  • The Verdict card. At test end, one number per campaign: true incremental ROAS versus platform-reported ROAS, with a confidence interval and a plain-English call — “Sponsored Listings on DoorDash returned $1.40 per dollar, not the $4.10 DoorDash reported. 66% of attributed orders were customers who already knew you.”
  • New-customer honesty check. Uses the platform’s own new-versus-existing customer split to separate genuine acquisition from paying to reach your regulars, and prices the actual cost of a new customer against their observed repeat value.
  • Visibility-collapse alarm. Watches organic (non-ad) order volume during holdout windows. If pausing ads genuinely craters your organic placement, the product tells you that too — and quantifies the ransom, so you know what you’re actually buying.
  • The reallocation memo. A one-page monthly PDF an operator can hand to a partner or a franchisor: what to cut, what to keep, what to test next, and the dollar impact.

6. AI angle — what’s load-bearing

Two places, and both are doing real work.

Forecasting the counterfactual. The core deliverable is “what would sales have been without the ads.” That’s a time-series problem with messy inputs — weather, holidays, local events, menu changes, competitor promotions, a driver shortage on a rainy Tuesday. Off-the-shelf forecasting gets you a baseline; the model earns its keep by learning each store’s idiosyncratic shape from limited history and honestly reporting when the interval is too wide to call a verdict. Remove this and there’s no product — you’re left comparing raw week-over-week numbers, which is exactly the naive analysis that produces wrong answers and lost trust.

Reading the mess. Every platform report is a differently-shaped CSV, the schemas change without notice, and campaign naming is whatever the operator typed at 1am. An LLM layer normalises these into a common event log and, more importantly, writes the verdict memo in language an owner-operator acts on rather than a dashboard they ignore. The failure mode of analytics products for SMBs is not bad math — it’s math nobody reads.

What is not AI here, deliberately: the experiment design. That’s statistics, and it should be boring, pre-registered, and auditable. If a customer ever asks “why should I believe you,” the answer has to be a method they can check, not a model they can’t.

7. Localization angle (if any)

N/A — this is a global play with an English-first launch across US and UK, then Australia and Canada.

The reason not to force localization: the wedge is the platform, not the jurisdiction. DoorDash, Uber Eats, Deliveroo (now inside DoorDash’s global commerce media platform) and Just Eat run the same auction mechanics and the same self-reported attribution everywhere. A restaurant in Manchester has the identical problem to one in Austin, and the report exports are structurally similar. Adding a language before you’ve proven the method in one market is a distraction.

One deliberate exception worth flagging for later: India. Swiggy and Zomato take an effective 25–35% all-in, and the pain there is sharper and uglier — multiple restaurateurs have alleged the platforms ran ads and discounts without consent, with unauthorised deductions reported from ₹10,000 to over ₹20 lakh, and India’s CCI investigation arm has found anti-competitive conduct including preferential treatment of some restaurant partners (Medianama, Outlook Business, Tribune). But that is a consent and recovery product, not an incrementality product — closer to this catalogue’s existing PayoutSnap than to LiftReferee. Different buyer psychology, different build. Don’t chase both.

8. Business model — path to $1M–$5M ARR

Pricing:

  • Solo — $149/mo, one location, both platforms, one active test at a time.
  • Multi — $99/location/mo, 2–20 locations, cross-location control groups (this is the good tier — more locations means better experiments).
  • Audit — $1,500 one-time, 90-day retrospective incrementality audit with no subscription. This is the door-opener, and it converts.

Benchmark sanity check: POS-attached marketing suites like Toast’s already charge $185–$379/month before a single campaign launches, and single-location operators carry $300–$800/month in barely-used subscriptions (Restaurant Velocity). A $149/mo tool that adjudicates $3,000/mo of spend is not a hard sell — it’s a rounding error against the thing it audits.

ACV: ~$1,800 for a single location, ~$5,000–$12,000 for a 4–10 location group. Blended target ACV $2,800.

Math to $1M ARR: 360 accounts at $2,800 blended. Realistically ~250 single-location at $1,800 ($450K) plus ~110 small groups at $5,000 ($550K). Achievable inside 18 months with the outbound motion in section 9.

Math to $5M ARR: 1,800 accounts blended, which requires two things to be true: (a) the audit-to-subscription conversion holds above 35%, and (b) at least one POS or aggregator partnership (Toast marketplace, Chowly, ItsaCheckmate) delivers distribution at low CAC. Without a channel partnership this is a $2–3M business, which — per the operator’s own filter — is a fine outcome, not a failure.

Expansion path: more locations, then more channels (Grubhub, Deliveroo, Just Eat, Instacart for the convenience-store adjacency), then the same method applied to Google Local Ads and Meta — because once an operator trusts your holdout verdict on delivery, “which of my other channels are lying to me” is the obvious next question. That’s the real ACV ladder: from delivery-ads referee to the operator’s whole marketing-truth layer.

9. Go-to-market wedge — first 100 customers

1. The public teardown, done on real numbers. Recruit 10 operators to run a free 90-day audit in exchange for anonymised publication. Publish each as a hard-numbers post: “This 4-location burrito chain was told $4.10 ROAS. It was $1.40. Here’s the test design.” Restaurant operators forward this stuff. This is the content engine and the credibility proof simultaneously — and because the underlying claim is contrarian and checkable, it earns coverage in the operator press (Restaurant Dive, QSR, Nation’s Restaurant News) that a generic analytics tool never gets.

2. Cold outbound to the visibly-advertising. Sponsored listings are public. Anyone can open DoorDash or Uber Eats in a given zip code and see exactly which restaurants are paying for placement, in which dayparts, in which categories. Build a list of 3,000 independents and small chains observed running sponsored placements across 30 US metros. Pitch: a screenshot of their own sponsored tile, plus one line — “you’re paying for this slot. Want to know if the orders were new?” Expect 2–4% to take a free audit, 30–40% of those to convert to paid. That’s 20–45 paying customers from the first list.

3. Accountants and restaurant CFO-for-hire firms. Firms like Harmony CPA publish on third-party delivery economics; fractional restaurant CFOs and bookkeeping shops each carry 20–80 restaurant clients and are already the ones explaining why the P&L looks wrong. Sign 15 as referral partners on 20% recurring. One good firm delivers 10 accounts. This channel is slower to start and then compounds — and it’s the one that survives when the cold-email conversion decays.

4. The delivery-management adjacency. Consultants who manage delivery platforms for restaurants exist and charge retainers. They are not competitors — they’re distribution. Offer white-label audits at $600 wholesale so the consultant resells at $1,500. They get a defensible deliverable; you get their client list.

5. Operator communities. r/restaurateur and r/KitchenConfidential, plus the Restaurant Owners & Operators Facebook groups and the state restaurant-association newsletters. Not a broadcast — show up with the free audit offer attached to a real published teardown. One good thread in a 100K-member operator community produces 20+ audit requests.

10. Build complexity — justification

Medium. The ingestion side is off-the-shelf: CSV normalisation from two-to-four report formats, optional Toast/Square POS OAuth for channel sales. The forecasting is standard time-series work with weather and holiday regressors — well-trodden libraries, no research required. The genuinely careful part is the experiment design layer: pre-registration, power calculation (telling a small operator honestly that their volume is too low for a 3-week test to detect anything is a feature), and refusing to publish a verdict when the interval is too wide. That discipline is a design problem more than an engineering one.

Two people, 10–14 weeks to a v1 that can run a real test end-to-end. The main schedule risk is not code — it’s that every test takes 3–6 weeks of calendar time to produce a result, so the feedback loop on product quality is slow. Budget for that: the first 10 pilot audits should start in week 4, not week 12.

11. Gating checklist

GatePass?Note
Legal in target market✅Customer uploads their own merchant reports and pauses their own campaigns. No scraping of platform internals, no ToS-violating automation. Read-only analysis of data the operator already owns.
Ethical — no harm / dark patterns✅The product’s whole value is telling a small business to stop spending money. The ethical risk is the inverse of the usual one — the discipline required is refusing to overclaim statistical certainty.
Market exists (evidence above)✅$4B+ annualised delivery ad revenue, 400,000+ advertisers on DoorDash alone, documented cases of ad spend exceeding commission, incumbents explicitly floored at $50K/mo spend.
1–5 person team can build this✅Two people, 10–14 weeks. Standard stack.
Launchable with <$50K / ₹40L✅Under $15K: infra, data subscriptions for weather, and the founders’ time.

All five pass.

12. Feasibility score

AxisWeightScoreNotes
Problem intensity2016/20Ad spend at 8–18% of gross sales on 3–6% net margins is genuinely painful, and it’s felt monthly when the payout statement lands. Not a 19 because it’s a slow bleed — nobody gets fined, no deadline forces action, and an operator can ignore it for another quarter. Deferred pain converts worse than a hard stop.
Demand evidence1513/15Multiple independent signals: documented spend levels, $4B+ ad revenue growing 73% YoY, an incumbent category (Otter, Revly, Chowly) already selling adjacent tooling to exactly this buyer, and enterprise vendors publicly refusing the segment. Docked 2 because I have industry-analysis evidence of the measurement gap rather than a pile of verbatim operator complaints specifically about incrementality — operators complain about fees, and I’m inferring the measurement complaint sits underneath.
Build feasibility1511/15Off-the-shelf ingestion and forecasting; the experiment layer needs real statistical care. 10–14 weeks for a pair. The slow calendar of test results is the hidden cost.
Distribution clarity1512/15The observable-advertiser list is the strongest asset here — you can literally see who your customers are by opening the app. Teardown content is proven in this vertical. Docked because cold email to restaurant owners has brutal open rates and the accountant channel takes months to warm.
Revenue mechanics1511/15Pricing is benchmarked against a real $185–$379/mo POS marketing suite and audits $3K/mo of spend, so willingness-to-pay is sound. The risk is churn: a customer who runs one audit, cuts 40% of spend, and cancels. The subscription has to keep earning its place, and that’s an unproven assumption.
Time to first revenue108/10The $1,500 retrospective audit can be sold and delivered semi-manually in week 3–4 before the product is finished. Not a 10 because the proof takes a full test cycle.
Defensibility103/10This is the weak axis and I’m not dressing it up. The method is publishable — I’ve just published it. Otter or Chowly could bolt on holdout testing in a quarter if they wanted to. What protects you is that they mostly don’t want to: their business model is selling more ad management, and a tool that tells customers to spend less is a strategic contradiction for them. That’s a real but soft moat — a conflict-of-interest moat, not a technical one. It buys 12–18 months, plus whatever accumulated benchmark data (“median true ROAS across 800 independents”) is worth by then.
Total10074/100

13. Qualitative modifiers

Founder-fit tags

technical-heavy · content-heavy

You need someone who can do honest statistics and someone who can write the teardowns that sell it. Those can be the same person. This is not a sales-heavy business — the teardown content is the sales motion, and it only works if the numbers underneath are real.

Key assumptions to validate (3–5)

  1. Assumption: Platform-reported ROAS materially overstates incremental ROAS for independents — the gap is large enough to matter, not 10%. How to test: Run 10 free audits before writing a line of production code. If median true ROAS comes in above 70% of reported, there’s no product here. I’d expect 30–50%; if it’s 85%, kill it.

  2. Assumption: Operators will actually pause campaigns on a schedule you hand them. The entire method depends on customer compliance with an experiment. How to test: In the 10 pilots, measure adherence. If fewer than 6 of 10 follow the pause schedule without hand-holding, the product needs to become services-heavy — which changes the business.

  3. Assumption: The verdict doesn’t cause immediate churn. Customer learns their ads are wasteful, cuts spend, cancels. How to test: Track pilot cohort at day 90 and day 180. If retention past the first verdict is under 60%, reprice around the audit as the primary product with a thin monitoring subscription behind it.

  4. Assumption: Pausing ads doesn’t collapse organic placement so severely that the test is unrunnable (or the answer is always “you must pay the ransom”). How to test: Measure organic order volume during pilot holdout windows. This is genuinely unknown and it’s the most interesting question in the whole idea — if the answer is “organic dies without ads,” that’s a different but equally sellable product: quantifying the ransom.

  5. Assumption: Report formats stay stable enough that ingestion isn’t a treadmill. How to test: Track schema changes across the pilot period. Budget maintenance accordingly.

Risk flags

  1. Platform dependency: Everything rests on merchant-facing CSV exports staying available. DoorDash and Uber have no obligation to keep them, and a product that publicly demonstrates their ROAS is inflated is not a product they’ll want to help. Mitigation: POS-side data as a fallback baseline, and never automate anything that touches their systems in a way that violates merchant ToS.
  2. Incumbent response: Otter or Chowly adds a “true lift” tab. Likely eventually, unlikely soon — it contradicts their revenue model. Watch for it as a kill signal.
  3. Statistical credibility: One overclaimed verdict that an operator disproves destroys the brand. The product must refuse to answer when data is thin, which means telling some paying customers “we can’t tell you yet.” That’s commercially painful and non-negotiable.
  4. Churn-after-verdict: The structural risk in any audit-shaped business — you solve the problem and the customer leaves. The monitoring subscription and channel expansion exist to answer this, and neither is proven.
  5. Slow-bleed problem: No deadline, no fine, no forced action. Per this catalogue’s own hard-won lesson, deferred pain converts badly without a weekly wedge. The monthly payout statement is the closest thing to a recurring trigger — the product should land in the inbox the day it arrives.

14. Structured verdict

Score:                  74/100
Verdict:                GO
Confidence:             Medium
Best-fit builder:       Technical founder comfortable with time-series and experiment
                        design, paired with someone who can write credible operator-facing
                        teardowns. Restaurant-industry contacts help but aren't required —
                        the advertiser list is publicly observable.
Time to revenue:        3–4 weeks to first paid retrospective audit; 10–14 weeks to
                        productised subscription.
Capital to launch:      Under $15K (~₹13L)
Top 3 assumptions to validate first:
  1. Reported vs true ROAS gap is large — run 10 free audits, kill if median true
     ROAS exceeds 70% of platform-reported.
  2. Operators comply with a pause schedule — measure adherence across those same
     10 pilots, need 6+ of 10 unassisted.
  3. Retention survives the verdict — track pilot cohort to day 180, need >60%.
Kill criteria:
  - Abandon if median true incremental ROAS across the first 10 audits comes in
    above 70% of platform-reported ROAS (the gap isn't worth paying to measure).
  - Abandon if fewer than 6 of 10 pilot operators execute the holdout schedule
    without hands-on management (the method doesn't survive self-serve).
  - Abandon if either DoorDash or Uber Eats removes merchant-facing daily campaign
    CSV export and no POS-side substitute reconstructs the baseline.
  - Abandon if Otter, Chowly or Revly ships credible holdout-based incrementality
    to sub-$10K/mo advertisers before your v1.

15. Next step — 1-week validation sprint

The point of this week is to find out whether the gap between reported and real ROAS is big. Everything else is downstream of that one number.

  • Day 1–2: Recruit 6 operators running $2K+/mo in delivery ads. Source them by opening DoorDash and Uber Eats in three metros, listing the visibly-sponsored independents, and cold-emailing 150 of them with one offer: a free ad audit, no strings, you keep the report. Target 6 acceptances — this doubles as a live test of channel #2 from section 9.
  • Day 3–4: Collect their last 90 days of Ads Manager and Uber Eats Manager CSV exports. Build the baseline by hand in a spreadsheet — no product, no code. For each, compute a retrospective estimate of incremental lift using natural experiments already sitting in their history: weeks where budget capped out early, campaigns that ran out of funds mid-month, dayparts with no coverage. Every advertiser has accidental holdouts in their past. Find them.
  • Day 5: Compare estimated true ROAS against platform-reported ROAS for each of the 6. Then get on the phone and put the number in front of them.

Falsifiable outcome: GO if median estimated true ROAS is below 70% of platform-reported and at least 3 of the 6 operators, shown their own number, agree on the call to pay $1,500 for a full prospective audit. Anything less on either half and the idea doesn’t clear — the gap is either too small to sell or too abstract to sell, and both are fatal.

Interested in a detailed proposal?

Get a deep-dive with market research, competitive analysis, and implementation roadmap.

Contact us

info@startupbasket.ai