SB StartupBasket
All ideas
72 /100 GO Low complexity

ShipProof — AI-use dossier for software agencies

Answers the AI-governance section of a client's security questionnaire with evidence from your own repos, not a written promise.

— views
Evaluation Scores
72/100

GO

Overall Score

15
Problem
12
Demand
13
Build
12
Distrib.
11
Revenue
8
Time
1
Defense

ShipProof

1. One-liner

Answers the AI-governance section of a client’s security questionnaire with evidence from your own repos, not a written promise.

2. Trend signal — why now?

Three things happened in the last nine months that turned “do you use AI?” from small talk into a gate on a signed contract.

Enterprise procurement grew an AI section. In 2026, procurement teams added dedicated AI-governance modules to standard vendor risk assessments. CAIQ, SIG Lite, and most internal vendor-risk templates now include sections on model provenance, training-data rights, AI subprocessor transparency, and alignment with ISO 42001 and the NIST AI RMF. SOC 2 Type II is now the baseline; the AI governance layer sits on top of it. This is not a future obligation — it is on the questionnaire your client sends you this quarter.

Buyers started asking dev shops to show their work. The 2026 vendor-selection advice is explicit: audit a vendor’s own AI workflow and governance policy before signing, and ask prospective vendors to show how AI appears in their QA, project scoping, and code review process. The framing that matters: “What separates vendors from competitors is proof of competence delivered before the sale, not claimed after it.” Clients are no longer satisfied with a policy PDF.

Contracts started demanding retained evidence with a clock on it. The MSA clause set that spread through agency-land in 2026 includes an audit-rights clause: agencies maintain records for three years including tool category, dates, and content scope; clients may request annual compliance summaries and must receive provenance documentation within ten business days of a regulatory request. Numonic’s Q1 2026 agency survey found 73% of agency MSAs contain no AI-specific clauses at all — which is precisely why the remaining 27% are being rewritten right now, and why the redlines land on the agency.

The capture problem, meanwhile, got solved for free. Git AI (1.6k stars, v1.3.1 April 2026) links AI-authored lines to the agent, model, prompt, and session that produced them. Exceeds Ink records line-level multi-tool attribution as portable Git Notes. The raw signal exists in the repo. Nobody turns it into the document a procurement officer will accept.

That is the whole idea. The plumbing is commodity and open-source; the deliverable is missing.

Provenance:

  • Signal 1 (demand): Enterprise procurement teams added AI-governance modules to CAIQ / SIG Lite / internal vendor-risk templates in 2026, covering model provenance and AI subprocessor transparency — https://www.aetos-data.com/answers-insights/enterprise-security-ai-questionnaires — observed 2026-08-24
  • Signal 2 (feasibility): Git AI (v1.3.1, Apr 2026) tracks AI-authored lines with agent/model/prompt/session metadata as Git-native records; Exceeds Ink stores line-level attribution as portable Git Notes — https://github.com/git-ai-project/git-ai — observed 2026-08-24
  • Signal 3 (economic): Agency MSA clause sets now mandate 3-year record retention and provenance documentation within 10 business days; 73% of agency MSAs currently have no AI clauses (Numonic Agency Survey, Q1 2026) — https://www.numonic.ai/blog/ai-clauses-agency-contracts-2026 — observed 2026-08-24 Category: Tech-unlock (attribution metadata became capturable in 2026) + Underserved niche (procurement-facing evidence layer, ignored by both the capture tools and the compliance-template vendors)

3. The opportunity

There are two camps here and neither one serves the agency.

Camp one: the capture tools. Git AI, Exceeds Ink, AI Footprint, QAInsights/code-provenance. These are excellent and mostly free. They are built by engineers, for engineers, and their output is a Git Note or a CLI table. They answer “which lines did Claude write?” That is a developer question. No procurement officer has ever asked it, and no client-facing document falls out of them.

Camp two: the compliance-template vendors. Numonic, DigitalApplied, Agent Mode AI, a growing pile of law-adjacent content shops selling clause packs and checklists. They tell an agency what to promise — disclosure, metadata preservation, audit rights, three-year retention. They sell the obligation. They do not sell any means of meeting it. Agent Mode AI’s own guidance is telling: the recommended process is “a standard completion checklist” per SOW. A checklist. Manually filled in, per project, forever.

The gap between those camps is a document. The agency has signed a clause promising provenance documentation in ten business days, the raw attribution is sitting in Git Notes nobody reads, and when the client’s security team actually sends the questionnaire, a senior engineer loses two days assembling a Word file by hand — or, more commonly, the agency writes a vague paragraph and hopes.

What a focused team does 10× better: stand between the repo and the questionnaire. Read the attribution metadata the free tools already capture, plus the CI and PR record, and emit the artifact procurement actually accepts — per-project, per-deliverable, with the human review step evidenced rather than asserted.

The sharp version of the insight: the valuable claim is not “AI wrote 40% of this.” Every agency is terrified of that number and no client wants to buy it. The valuable claim is “every AI-authored line in this deliverable was reviewed and approved by a named human, and here is the record.” That reframes the number from a liability into a quality attestation. The 2026 research makes this the live worry — 78.9% of agentic PRs get single-developer review, agentic PRs have 5.3x longer pickup time, and CodeRabbit’s research puts 1.7x more issues in AI-generated PRs. Clients are right to ask. Agencies that can answer crisply will win work from agencies that can’t.

Honest scoping of the regulatory angle. I looked hard at EU AI Act Article 50 and California SB 942 because the template vendors lean on both, and the lean is oversold. Article 50’s obligations bite on deepfakes and public-interest text — source code is not covered. SB 942 (operative 2 Aug 2026, $5,000 per violation per day) binds generative-AI providers with >1M monthly California users, which is not a 12-person dev shop. The binding force here is contractual and commercial, not statutory. That is a weaker “why now” than a regulation with a deadline, and I’ve scored it accordingly. It is also, notably, more durable: procurement requirements don’t get repealed or extended.

4. Target market

  • Primary customer: Owner or delivery lead at a software agency / dev shop / product studio, 3–40 staff, billing €50–€500/hr (or $75–$250/hr US), selling into mid-market and enterprise clients that run a formal vendor-security process. Concentrated in EU/UK, US, and the India/Poland/LatAm outsourcing corridors selling into Western buyers.
  • Why they buy: Not for compliance virtue. They buy because an AI-governance section appeared on a questionnaire standing between them and a signed contract, and because a client’s counsel put a 10-business-day provenance clause in the MSA they already signed. The purchase moment is a deal at risk, which is the only moment agencies reliably spend money.
  • Rough TAM reasoning: This is deliberately a narrow slice. Globally there are on the order of tens of thousands of software agencies in the 3–40 band that sell to buyers with formal procurement. I don’t need a precise number — I need ~800 of them at ~$150/mo to clear $1.4M ARR. The constraint is not market size, it’s reaching them.
  • Why now for them: Small teams are the heaviest users of AI coding agents — teams with 1–5 contributors averaged 50.2 agentic PRs per repository over three months, far above medium and large teams. The agencies with the most AI-authored code in their deliverables are exactly the ones with the least governance apparatus. They are maximally exposed and minimally equipped.

5. Product sketch (MVP)

  • Repo connect. Point ShipProof at GitHub/GitLab repos. It ingests existing AI attribution (Git AI notes, Exceeds Ink notes, commit trailers, co-author tags) and, where none exists, installs the capture shim.
  • Review-evidence linking. For every AI-authored change, bind it to the PR that merged it, the human who approved it, the checks that ran, and the timestamp. This is the artifact that converts an embarrassing percentage into a defensible process claim.
  • Per-deliverable dossier. Select a client, project, or release tag → generate a dated PDF/HTML record: AI tools and model versions used, scope of AI-assisted work, human review attestation per change, and the vendor terms (zero-retention, commercial-use, training-data exclusion) for each tool in the stack.
  • Questionnaire autofill. Paste or upload the client’s AI-governance section (CAIQ/SIG-style). ShipProof drafts answers grounded in the agency’s actual repo evidence, flagging every question it cannot substantiate rather than bluffing.
  • Standing answer library. Answers approved once are reused across every future questionnaire, with staleness warnings when the underlying tool stack changes.
  • Tool inventory + terms watch. Tracks which AI coding tools the team uses and alerts when a vendor changes retention or training-data terms — the thing that silently invalidates last quarter’s attestation.
  • Retention vault. Three-year immutable record store, because that is what the clause requires and no agency wants to own that problem.

6. AI angle — what’s load-bearing

Two places, and I want to be precise because the weak version of this product is “a database with a PDF button.”

Load-bearing use #1 — questionnaire mapping. Every enterprise buyer’s AI-governance section is worded differently and none map cleanly to a fixed schema. Turning “Describe your controls governing the use of generative AI in the development of deliverables provided to Customer” into a grounded answer assembled from repo evidence is genuinely a language task. Doing it against a fixed template fails on question two.

Load-bearing use #2 — scope classification. Deciding whether a given AI-authored change is materially in scope for a client deliverable (versus internal tooling, tests, scaffolding, or a different client’s project) requires reading diffs and repo structure with judgment. Getting this wrong in either direction is fatal: over-report and you hand the client a scary number; under-report and you’ve made a false attestation in a contract.

Remove the AI and you have a metadata viewer that emits a form letter — which is roughly what the free capture tools already are. The AI is what makes the output client-ready instead of engineer-ready.

Where AI is explicitly not load-bearing: the attribution capture itself. That is deterministic metadata plumbing, it is solved, and it is open-source. Pretending otherwise would be exactly the “AI sticker on the side” the persona warns about.

7. Localization angle (if any)

N/A — this is a global play. The buyer is defined by who they sell to, not where they sit. A Kraków or Bengaluru dev shop selling into a German industrial client faces the identical questionnaire as a Manchester one. Language matters only in that questionnaires arrive in German, French, and Dutch — which is a feature (multilingual questionnaire ingest), not a localization strategy. No payment-rail or regulatory localization wedge exists here.

8. Business model — path to $1M–$5M ARR

  • Pricing: $99/mo (Solo, up to 3 devs, 5 repos) · $249/mo (Studio, up to 15 devs, unlimited repos, questionnaire autofill) · $599/mo (Agency, unlimited seats, retention vault, white-label dossiers for clients). Annual billing at 2 months free.
  • ACV: ~$2,400 blended, assuming the mix skews to Studio.
  • Rough math to $1M ARR: 420 customers × $249/mo × 12 ≈ $1.25M. At the low end, 840 Solo customers — but the real mix lands in between.
  • Rough math to $5M ARR: ~1,700 customers at blended ACV, which is a stretch through self-serve alone. The credible path is a second SKU: the client-side seat, sold to the enterprise buyer who wants a standing view across all their vendors. That flips the model from selling anxiety to selling oversight and roughly triples ACV. I’d treat this as a year-two hypothesis, not a plan.
  • Expansion path: Seats grow with headcount; retention vault and white-label dossiers are natural upsells; the questionnaire library becomes stickier every quarter as answers accumulate.

Gross margin is high — this is metadata processing plus modest LLM spend on questionnaire drafting, call it $6–12/customer/month at Studio volume.

9. Go-to-market wedge — first 100 customers

The lead magnet writes itself, and it’s the same trick that works in every compliance-adjacent category: show them they’d fail.

  • The free “AI Governance Readiness Scan.” Connect a repo, get a one-page score against the actual AI-governance questions on SIG Lite and CAIQ. Almost every agency fails at least half — not because they’re careless, but because nobody has ever asked them to evidence review-per-change before. That failing scorecard is the sales conversation. This is the exact playbook ClearRun-style diagnostics use and it converts because the result is specific and personally embarrassing.
  • Clutch / DesignRush / The Manifest scrape. These directories list software agencies with size bands and named principals. Filter to 3–40 staff shops tagged for enterprise clients, ~4,000 targets. Personalized outreach leading with the scan, not the product. Realistic: 25–30% open, 4–6% run the scan, 15–20% of scanners convert to trial. That’s ~30 trials per 1,000 contacted.
  • Ride the clause vendors. Numonic, Agent Mode AI, DigitalApplied and the employment/tech lawyers publishing these clause packs have exactly our audience and sell them an obligation with no fulfilment path. Affiliate or co-marketing deal: their clause pack ships with a ShipProof trial as “how you actually do clause 5.” Warm, high-intent, and they have no competing product.
  • r/agency, r/ExperiencedDevs, Indie Hackers, and the Clutch-adjacent agency-owner Slacks/Discords. Not “content marketing” — specifically, publish the annotated real AI-governance questionnaire sections with model answers. That artifact gets bookmarked and passed around by exactly the person who signs up.
  • Fractional CTO / agency-coach network. A small number of consultants advise dozens of agencies each on ops and sales process. Ten of these people reliably introduced is worth more than 1,000 cold emails.

The honest weak point: agencies are notoriously cheap and notoriously bad at buying tools until a deal is on fire. The scan exists to manufacture the fire.

10. Build complexity — justification

Low. The attribution capture is open-source and Git-native — we integrate rather than invent. The remaining surface is a Git host integration, a document generator, an LLM call for questionnaire mapping, and a retention store. A competent pair ships a credible v1 in 8–10 weeks; a solo builder in 12–14. The only genuinely fiddly parts are (a) handling repos with no prior attribution data, where we have to install capture and accept that history before install-day is unprovable, and (b) not over-claiming in generated attestations, which is a prompt-discipline and UX problem rather than an engineering one.

11. Gating checklist

GatePass?Note
Legal in target market✅Reads customer-owned repos under explicit authorization; produces documents the customer signs. No regulated activity.
Ethical — no harm / dark patterns✅One real hazard: this must not become a laundering tool for false attestations. Product must flag unsubstantiated answers rather than fill them — a design constraint, not an afterthought.
Market exists (evidence above)✅Procurement AI modules are live in CAIQ/SIG Lite; MSA clauses with 10-day provenance demands are in circulation.
1–5 person team can build this✅8–10 weeks for a pair on off-the-shelf components.
Launchable with <$50K / ₹40L✅Two people, three months, modest LLM spend. Well under.

All five pass.

12. Feasibility score

AxisWeightScoreNotes
Problem intensity2015/20Sharp and expensive at the moment it bites — a stalled contract — but episodic, not daily. Between questionnaires the pain vanishes, which caps this below the 17+ band.
Demand evidence1512/15Strong indirect evidence: procurement modules are real, clause packs are selling, capture tools have traction (1.6k stars). Weak spot — I found no agency saying “I would pay for this,” only third parties describing the obligation.
Build feasibility1513/15Capture layer is free and open-source; we build the document and mapping layer. Pair, 8–10 weeks.
Distribution clarity1512/15Named directories (Clutch/DesignRush), a specific and proven diagnostic lead magnet, and an unusually clean partner channel in the clause vendors. Docked for agencies’ well-earned reputation as slow, cheap buyers.
Revenue mechanics1511/15$99–599 is well inside agency tool budgets and $1M ARR needs only ~420 mid-tier customers. Docked because the $5M path genuinely depends on an unvalidated second SKU.
Time to first revenue108/10Scan → trial → paid is short, and the trigger event (a live questionnaire) creates urgency. Realistically 6–8 weeks post-launch.
Defensibility101/10This is the problem. The capture layer is open-source, the document layer is a weekend for a competent team, and GitHub could ship “AI provenance reports” as a free Advanced Security feature and end the category overnight. The only accruing asset is the questionnaire-answer library. Being honest: there is essentially no moat here.
Total10072/100

13. Qualitative modifiers

Founder-fit tags

technical-heavy · sales-heavy

Technical to build the Git and attribution integrations credibly; sales-heavy because agencies do not buy compliance tooling passively and the diagnostic-to-deal motion is hands-on. A builder who won’t do outbound will watch this die with a great product.

Key assumptions to validate (3–5)

  1. Assumption: Agencies are actually being asked AI-governance questions frequently enough to feel it — not just once a year. How to test: Ask 30 agency owners for their last three client security questionnaires. Count how many contain an AI section and how long each took to answer. If under a third do, the trigger is too rare and pricing collapses.
  2. Assumption: The buyer will pay to answer faster, rather than continuing to wing it with a vague paragraph that clients accept anyway. How to test: Present two agencies’ real questionnaire responses to five enterprise security reviewers; ask whether the vague answer would pass their review. If it passes, there is no product.
  3. Assumption: The “reviewed by a named human” reframe is genuinely more sellable than the raw AI percentage. How to test: A/B two versions of the dossier with 15 agency owners; see which one they’d willingly hand a client. If both scare them, the artifact is wrong.
  4. Assumption: GitHub/GitLab won’t ship this natively within 12 months. How to test: Track GitHub Advanced Security and Copilot enterprise roadmaps and changelogs monthly. Unfalsifiable in advance — hence a monitoring commitment, not a one-time check.

Risk flags

  1. Platform obsolescence (severe). GitHub shipping native AI provenance reporting is the single most likely killer. The mitigation is to be the multi-tool, multi-host, client-facing document layer GitHub has little incentive to build — but that is a thin hedge, not a moat.
  2. Attestation liability. If ShipProof generates a dossier an agency signs and it turns out materially wrong, we are adjacent to the resulting dispute. Requires careful product framing (we assemble evidence; the agency attests) and a hard rule against auto-filling unsubstantiated answers.
  3. Trigger frequency risk. If the AI-governance section turns out to be a once-a-year event per agency, this is a $200 one-off report, not a subscription. This is the assumption most likely to be wrong and it is scored into demand and revenue above.
  4. Regulatory framing is softer than it looks. SB 942 doesn’t bind these buyers and Article 50 doesn’t cover code. Anyone marketing this as “AI Act compliance” is overreaching, and a client’s counsel will catch it. Sell the procurement reality, not a fake statutory deadline.
  5. Blank-history problem. Agencies that install capture today cannot evidence anything they shipped before today. For a client asking about a project delivered last year, ShipProof has nothing to say — an awkward gap in exactly the first sales conversation.

14. Structured verdict

Score:                  72/100
Verdict:                GO
Confidence:             Medium
Best-fit builder:       Technical founder who has run or sold into an agency, comfortable with
                        outbound; ideally has personally filled out a client security questionnaire
Time to revenue:        6–8 weeks post-launch (10–12 weeks from zero)
Capital to launch:      $8–15K / ₹7–13L
Top 3 assumptions to validate first:
  1. Frequency — collect the last 3 client security questionnaires from 30 agency owners;
     confirm >1/3 contain an AI-governance section
  2. Sufficiency — show 5 enterprise security reviewers a typical vague AI answer;
     confirm they would reject it
  3. Artifact fit — A/B the "review-attested" dossier vs. raw AI-percentage report with
     15 agency owners; confirm the former is one they'd hand a client
Kill criteria:
  - Abandon if <30% of 30 surveyed agencies have seen an AI-governance section in the last 6 months
  - Abandon if GitHub ships native per-deliverable AI provenance reporting before v1 launch
  - Abandon if <5 of the first 40 readiness scans convert to a paid trial

The score is capped by a 1/10 on defensibility and I’m not going to dress that up. This is an execution-and-speed play in a window that may close when a platform vendor notices. The reason it’s still a GO rather than a VALIDATE: the build is cheap, the trigger event is commercially urgent, the distribution channel is unusually concrete, and the partner channel (clause vendors selling an obligation with no fulfilment) is a genuine unfair advantage for whoever moves first.

15. Next step — 1-week validation sprint

  • Day 1–2: Pull 400 agencies off Clutch in the 3–40 staff band tagged for enterprise clients. Email 120 with one question: “Has a client sent you a security questionnaire with an AI section in the last 6 months? Reply Y/N and I’ll send you the annotated question bank.” The question bank is real and worth having, so the reply rate measures something honest.
  • Day 3–4: Book 10 calls with responders. On each call, ask them to screen-share how they answered it last time. Time how long it took. Ask what they’d pay to have it answered in an hour. Separately, take one real questionnaire’s AI section to 5 enterprise security reviewers and ask whether a typical agency answer would clear their review.
  • Day 5: Hand-build the dossier for one willing agency’s repo, entirely manually. Watch their face. Then ask them to send it to an actual client.

Falsifiable outcome: Go only if ≥35% of the 120 answer “yes, in the last 6 months,” AND ≥4 of 10 call subjects name a number ≥$99/mo unprompted, AND ≥3 of 5 security reviewers say the typical vague answer would fail their review. Miss any one of the three and this is a report, not a subscription — and I’d rather learn that in week one than month six.

Interested in a detailed proposal?

Get a deep-dive with market research, competitive analysis, and implementation roadmap.

Contact us

info@startupbasket.ai