SB StartupBasket
All ideas
76 /100 GO Low complexity

CaseScrub — recordability referee for high-hazard employers

Decides which injuries go on your OSHA log, then strips names from the narratives before they go public.

— views
Evaluation Scores
76/100

GO

Overall Score

16
Problem
12
Demand
13
Build
11
Distrib.
11
Revenue
8
Time
5
Defense

CaseScrub

1. One-liner

Decides which injuries go on your OSHA log, then strips names from the narratives before they go public.

2. Trend signal — why now?

OSHA’s July 2023 final rule rewired what a mid-size employer owes the federal government every March. It’s not the old one-page summary anymore.

Under the expanded 29 CFR 1904.41, an establishment with 100 or more employees in one of roughly 100 high-hazard NAICS codes (Appendix B) must electronically submit Form 300 and Form 301 case-level data — not just the 300A summary. That means the date, physical location and severity of each injury, details about the worker, and free-text narrative describing how the incident occurred. Holland & Knight put the reach at more than 52,000 employers.

Three things make this urgent rather than routine:

The data goes public. OSHA publishes the submitted case data for public access. Holland & Knight explicitly flags that public disclosure “may negatively impact company reputation based on injury records.”

OSHA scrubs the structured fields — not your prose. When you submit 301 data through the ITA, OSHA excludes Employee Name, Employee Address, Physician Name and Facility Address, and applies automated plus manual review to catch names and SSNs. But fields 14–17 are free text and the duty to sanitize them sits with the employer. OSHA’s guidance tells establishments to review the log and “remove non-mandatory information that could reasonably be expected to directly identify individuals” before submitting. Holland & Knight goes further: even “gender, race, ethnicity or occupation” can indirectly identify a worker at a single-site employer. Nobody is doing this scrub for you.

Missing the deadline draws an inspection, not just a fine. The CY2025 deadline was March 2, 2026. Non-submission is an Other-Than-Serious citation with a statutory maximum of $16,550, and OSHA has a six-month citation window. Worse: non-submitters get dropped into the “non-responder pool” under the Site-Specific Targeting program, from which OSHA randomly selects facilities for programmed inspections. You don’t just pay — you get visited.

And the underlying judgment call is the one everyone gets wrong. Determining “medical treatment beyond first aid” is the single most common recordkeeping error. Over-record and you distort your own DART/TRIR and hand a public dataset a worse safety record than you earned. Under-record and, per the guidance, that’s where most OSHA recordkeeping penalties come from.

Provenance:

3. The opportunity

The EHS software market sells safety programs. This customer doesn’t have a safety program — they have a deadline and a liability.

Look at what’s actually on offer. SafetyCulture runs $24–36/user/month. SmartQHSE is $9/user/month across broader HSE workflows. VelocityEHS starts around $1,000–2,000/month. Cority is $40,000+/year. Mid-market employers pay $3,000–15,000/year for full 300/300A/301 + ITA across 5–25 establishments. These are per-seat platforms priced for companies staffing a safety department.

The newly-covered establishment is a single site with 100–400 employees in a metal fab shop, a plastics plant, a food processor, a nursing facility, a warehouse. Their “safety department” is an HR generalist or an ops manager wearing a second hat. They were exempt from case-level submission until this rule caught them. Buying a per-seat EHS suite to satisfy one annual filing is absurd — it’s a $9,000/year answer to a once-a-year problem.

And here’s what none of those suites sell as a headline capability: narrative redaction. They will store your 301s and push them to the ITA. They will not read “Jose slipped on the ramp by the east bay while the second-shift lead was calling him over” and tell you that at a 140-person single-site employer, that sentence identifies a specific human being to anyone who works there — and that you are about to publish it on a federal website.

The wedge is the two acts of judgment the incumbents leave to the customer:

  1. Is this case even recordable? (the first-aid vs. medical-treatment gray zone)
  2. Does this narrative expose a worker once it’s public?

Both are language-and-rules problems. Both are exactly what a good model does well, and exactly what a form-storage platform does not do at all.

4. Target market

  • Primary customer: The EHS-responsible manager — often titled Safety Manager, HR Manager, or Operations Manager — at a single-establishment employer with 100–400 employees in an Appendix B high-hazard NAICS code. Manufacturing, food processing, warehousing, nursing and residential care, waste management, building services. Typically no full-time safety staff; often one person who inherited the OSHA log.
  • Why they buy: They are newly in scope for case-level reporting, the data becomes public, they have no confident answer to “is this recordable,” and the downside of guessing wrong is a citation up to $16,550 plus landing in a pool that feeds programmed inspections. The felt pain is concentrated in January–February but the recordability decision happens every time someone gets hurt, all year.
  • Rough TAM reasoning: Holland & Knight sizes the rule’s reach at more than 52,000 employers. Trim to single-site, 100–400 employee shops without a dedicated EHS platform and a serviceable base of roughly 15,000–25,000 establishments is defensible. At $1,200/year that’s a $18–30M addressable line — too small for the EHS incumbents to re-architect around, more than large enough for a sub-$5M ARR business.
  • Why now for them: The rule newly captured them. Before this, a 140-person plastics plant filed a one-page 300A summary that nobody read closely. Now they ship every case narrative into a public database annually, and OSHA is running weekly matches between open inspections and employers that failed to submit.

5. Product sketch (MVP)

  • Recordability referee — describe the incident in plain language (“cut his hand, urgent care used two stitches and gave a tetanus shot”), get a recordable / not-recordable call with the specific 1904.7 criterion cited, plus the counter-argument so the manager can overrule it knowingly.
  • Narrative PII scan — reads every Form 301 free-text field (14–17) and flags names, nicknames, job titles, and the indirect identifiers Holland & Knight warns about — gender, race, ethnicity, occupation — with a one-click rewrite that preserves cause and severity as 1904.29 requires.
  • Single-site identifiability check — the differentiator. Given your headcount and department structure, tells you when a scrubbed narrative still fingers one person (“the only forklift operator on third shift”) and needs further generalization.
  • 300 / 300A / 301 log — clean, boring, one establishment, no seat licenses. Enough to hold a year of cases and print the 300A for the February 1 posting.
  • ITA-ready export — produces the CSV in the exact ITA submission shape, so filing is upload-not-retype.
  • Deadline clock — February 1 posting, March 2 submission, and the five-year retention rule, with escalating reminders.
  • Decision trail — every recordability call, who made it, what the reasoning was, and what changed. This is what you hand an inspector when they ask why case #7 isn’t on the log.

6. AI angle — what’s load-bearing

Remove the model and this collapses into a spreadsheet with a due date — which is what these employers already have and already get wrong.

Two jobs are irreducibly language work:

Recordability classification. The input is an unstructured human account of an injury plus a clinic note. The rule is a fact-specific standard with a long tail of OSHA letters of interpretation — Active Release Techniques, stretching, and musculoskeletal treatment all have their own guidance. Mapping messy prose onto 1904.7 with the right criterion cited is a reasoning task. A decision-tree UI has been tried for twenty years; people still get it wrong, which is why it remains the most common recordkeeping error.

Identifiability judgment. Regex finds “Jose.” Regex does not find “the second-shift lead,” “the new guy on the press brake,” or “she.” Deciding whether a sentence re-identifies a worker inside a 140-person plant requires holding the narrative and the org context together and reasoning about who could be picked out. That is a language model’s home turf and nothing else does it.

The AI is doing the thing the customer is actually afraid of getting wrong. That’s load-bearing.

7. Localization angle

N/A — this is a US-only play, and deliberately so. The product is a wrapper around 29 CFR 1904 and the ITA submission format. That specificity is the moat; there is nothing to localize. The closest adjacency is the ~22 OSHA State Plan states that operate their own programs with their own quirks, which is a within-US expansion path rather than a geographic one.

8. Business model — path to $1M–$5M ARR

  • Pricing: $99/month per establishment, billed annually at $1,188/year. One flat price, unlimited users. The pitch is explicitly anti-per-seat: your ops manager, HR generalist and plant supervisor all get logins because injuries get reported by whoever is standing there.
  • Second tier: $249/month for multi-establishment employers (2–10 sites), which is where the accidental expansion happens as customers realize a sister plant is also in scope.
  • ACV: ~$1,200 blended, drifting to ~$1,500 as multi-site accounts mix in.
  • Rough math to $1M ARR: 830 establishments × $1,188 = $986K. Call it 850 customers. Against a serviceable base of 15,000–25,000, that’s 3–5% penetration.
  • Rough math to $5M ARR: ~3,300 single-site accounts, or more realistically ~2,000 single-site plus ~400 multi-site accounts at $2,988. Requires expanding beyond Appendix B into the 20–249 employee 300A-only tier, where the recordability question is identical even though the submission burden is lighter. That tier is far larger and gets served by the same product with the 301 module switched off.
  • Expansion path: additional establishments → State Plan variants → workers’ comp claim-linkage (the recordability call and the comp claim are triggered by the same event and are currently decided by different people from different files).
  • Margin: inference on a few hundred incidents per customer per year is rounding-error cost. This is a >90% gross margin product.

9. Go-to-market wedge — first 100 customers

The distribution advantage here is unusual and worth being precise about: OSHA publishes the submitted data, so the target list is a public federal dataset.

  • Mine the ITA public data for the non-responder-shaped list. OSHA publishes establishment-level submitted data. Filter to Appendix B NAICS codes, 100–400 employees, single establishment. This yields a named list with address and industry — thousands of exactly-qualified targets. Cross-reference year-over-year: an establishment that submitted in a prior year and went missing is a live compliance problem. Direct mail plus email, referencing their own published injury summary. Expect 2–4% to a paid trial on a list this well-qualified.
  • February panic channel. Buy search on “OSHA ITA submission,” “OSHA 300 log deadline,” “is this OSHA recordable” in January–February. This is intent traffic with a hard date attached and low commercial competition versus general EHS terms. The seasonality is real and it’s a feature: land them in panic season, keep them because recordability decisions happen year-round.
  • Workers’ comp brokers and TPAs as a channel. Brokers serving 100–400 employee manufacturers already advise on injury classification because it moves the client’s experience mod — M3 Insurance publishes a first-aid-vs-medical-treatment reference guide, which tells you brokers are fielding this question for free. Recruit 10–15 regional brokers with a revenue share; each carries 30–100 in-scope clients and an existing trusted channel.
  • The recordability answer engine as content. “Is a tetanus shot recordable? Is Active Release Technique first aid?” — every one of these is a real search with a real OSHA interpretation letter behind it. Build the free public answer tool, gate the log behind signup. This ranks because the queries are specific and the correct answers are genuinely hard to assemble.
  • Trade association safety committees. Regional manufacturers’ associations and food processor groups run safety committees where this rule is already an agenda item. A 20-minute “what the 300/301 rule actually requires now” session is a welcome slot, not a sales pitch.

10. Build complexity — justification

Low. Standard web stack, a model API, and one CSV export format. The 300/300A/301 forms are fixed and published; the ITA submission shape is documented in OSHA’s ITA User Guide. No integrations required for v1 — the customer uploads or types incidents, and downloads a file to upload to OSHA. There is no API to certify against and no approval to obtain.

The real work is not engineering, it’s domain encoding: assembling 1904.7 and the OSHA letters of interpretation into a reliable classification behavior, and tuning the identifiability check so it doesn’t cry wolf on every narrative. Budget 8–10 weeks for a solo technical builder, with a certified safety professional on retainer to adjudicate the classification test set. That retainer is the single most important line item and should not be skipped.

11. Gating checklist

GatePass?Note
Legal in target market✅Recordkeeping assistance software. Not legal advice; ships with the customer as final decision-maker and a documented override trail.
Ethical — no harm / dark patterns✅Net effect is more accurate injury records and less worker PII in a public database. Explicitly refuses to help suppress genuinely recordable cases — see risk flags.
Market exists (evidence above)✅52,000+ employers newly in scope; incumbents charging $3,000–15,000/yr; $16,550 penalty exposure.
1–5 person team can build this✅Solo technical builder plus a part-time CSP advisor.
Launchable with <$50K / ₹40L✅Realistically $12–18K: advisor retainer, inference, direct mail for the first list.

All five pass.

12. Feasibility score

AxisWeightScoreNotes
Problem intensity2016/20Real teeth — $16,550 plus programmed-inspection exposure, and public reputational data. Docked because the acute pain is seasonal (Jan–Mar); the year-round recordability pain is real but quieter.
Demand evidence1512/15Strong structural evidence: 52,000+ employers, priced incumbents, documented penalty regime, documented “most common error.” Docked for thin direct customer voice — I could not surface verbatim operator complaints, only practitioner-guidance content.
Build feasibility1513/15Fixed published forms, one export format, no integrations, no API certification. Domain encoding is the only hard part.
Distribution clarity1511/15Unusually good target list — OSHA publishes the establishment data, so the prospect list is public and precisely filterable. Docked because cold outreach to non-technical ops managers converts slowly and the peak window is narrow.
Revenue mechanics1511/15$1,188 ACV against $3,000–15,000 incumbents is a clean value story, and 850 customers for $1M ARR is achievable. Docked: annual-event products carry real churn risk once the filing is done.
Time to first revenue108/10Sellable pre-build against the March deadline; a January launch could pre-sell. Off-cycle launches wait longer.
Defensibility105/10Honest assessment: an EHS incumbent could ship narrative scrubbing as a feature. The defense is the accumulated classification test set, niche brand, and the fact that this segment is too small for incumbents to re-price around. Execution moat, not a structural one.
Total10076/100

13. Qualitative modifiers

Founder-fit tags

technical-heavy · domain-expertise-required

The domain tag is not optional. Shipping a recordability classifier without a certified safety professional validating the test set is how you get a customer cited and a lawsuit. Budget for the advisor from week one.

Key assumptions to validate (3–5)

  1. Assumption: Newly-covered single-site employers will pay ~$1,200/year for a filing-plus-judgment tool rather than absorbing it into an existing EHS seat or their broker’s free advice. How to test: 40 calls to safety/HR managers at Appendix B establishments pulled from the public ITA data; ask what they used for the March 2026 filing and what it cost them in hours.
  2. Assumption: Narrative PII exposure is a felt fear, not just a lawyer’s theoretical concern. How to test: pull actual published 301 narratives from the ITA public dataset, find real examples that identify workers, and show them to 20 target managers. Measure whether “that could be us” lands.
  3. Assumption: The recordability classifier can hit ≥90% agreement with a certified safety professional on a blind set. How to test: build a 200-case test set from OSHA letters of interpretation and real scenarios; have the CSP grade it blind before writing any marketing.
  4. Assumption: Churn after the March filing is survivable. How to test: structure the pilot as annual-prepay and track whether customers log incidents in June–September. Year-round logging is the retention signal.

Risk flags

  1. Seasonality / churn: This is the central commercial risk. A once-a-year deadline product invites “file and cancel.” Mitigation is making the year-round recordability call the daily habit, and annual prepay. If usage flatlines after March, the business is a service, not a SaaS.
  2. Incumbent feature response: Narrative scrubbing is a plausible quarter’s work for SafetyCulture or SmartQHSE. The bet is that they won’t re-price to $99/establishment for a segment that doesn’t fit per-seat economics. That bet could be wrong.
  3. Liability and misuse: A tool that decides recordability will attract customers hoping it says “not recordable.” The product must refuse to be a suppression instrument — default to the conservative call, show the counter-argument, log every override. Under-recording is where OSHA penalties concentrate; a product that helps under-record is both unethical and a lawsuit magnet.
  4. Regulatory reversal: The electronic submission rule has been expanded and contracted across administrations — the 2018 rule rolled back parts of the 2016 version. Another rollback would shrink the 301 submission burden, though the recordability and 300-log obligations survive any rollback.
  5. Thin direct customer voice: I could not source verbatim complaints from safety managers, only practitioner guidance and vendor content. The demand case rests on structural evidence rather than heard pain. That’s the weakest link in this proposal and it’s what the validation sprint must attack first.

14. Structured verdict

Score:                  76/100
Verdict:                GO
Confidence:             Medium
Best-fit builder:       Solo technical founder with a certified safety professional (CSP) on retainer
Time to revenue:        6–10 weeks (faster if launched into the Jan–Feb deadline window)
Capital to launch:      $12–18K (advisor retainer, inference, first direct-mail list)
Top 3 assumptions to validate first:
  1. Willingness to pay ~$1,200/yr — 40 calls to safety/HR managers at Appendix B establishments from the public ITA list
  2. Narrative PII exposure is felt, not theoretical — show real published 301 narratives that identify workers to 20 target managers
  3. Classifier hits ≥90% blind agreement with a CSP on a 200-case test set built from OSHA interpretation letters
Kill criteria:
  - Abandon if fewer than 6 of 40 called managers say they spent more than 4 hours on the March 2026 filing
  - Abandon if the classifier cannot clear 90% blind agreement with a CSP after two tuning rounds
  - Abandon if fewer than 30% of pilot customers log any incident in the 90 days after their March filing (confirms this is a service, not SaaS)

15. Next step — 1-week validation sprint

  • Day 1–2: Pull the public ITA establishment dataset. Filter to Appendix B NAICS, 100–400 employees, single site. Build the real target list and count it — this alone confirms or kills the TAM assumption with federal data rather than estimation. Simultaneously pull published 301 narratives and hunt for real examples where a worker is identifiable. Those examples are the entire sales pitch; if they don’t exist, the PII angle is dead.
  • Day 3–4: Call 40 managers off that list. One question set: what did you use for the March filing, how many hours did it take, who decided whether each case was recordable, and did anyone review the narratives before they went to OSHA. Do not pitch. Listen for whether they knew the narratives become public.
  • Day 5: Build the 200-case classification test set from OSHA letters of interpretation and have a CSP grade a model’s blind output.

Go/no-go: proceed only if (a) the filtered list exceeds 10,000 establishments, (b) at least 15 of 40 managers report spending 4+ hours on the filing or express surprise that narratives are public, and (c) blind classification agreement clears 90%. Any one of those failing means fix it or drop it — all three are measurable within the week, and none of them depend on my opinion.

Interested in a detailed proposal?

Get a deep-dive with market research, competitive analysis, and implementation roadmap.

Contact us

info@startupbasket.ai