Strategic thesis
decay: 12moThe argument in five bullets, before any tables. If only one section survives a rewrite, this is the one.
- The wedge is the audit itself. A 10-minute, free, AI-run audit that returns a ranked, dollar-quantified, builder-ready list of automations for one business — plus an explicit "don't automate" list. Build/maintain monetizes; the audit hooks. Comparable audits from McKinsey-style firms cost 8 weeks and six figures.
- We compete on specificity, not capability. The market is saturated with "AI for ops." Our defensible asset is the 200-row, citation-backed process database (§3) plus the 50-trap anti-list. Both signal we know which automations actually ship.
- The opening move is the never-automated tier. 30+ processes in our database are flagged
never_automated_yet— high pain, low awareness, no competitive density. Examples: vendor-renewal alerter, scope-creep tracker on services contracts, customer-health rollup from telemetry+tickets, partner-attribution stitcher. We don't win the first inbound-lead-routing deal; we win the "I didn't know that existed" deal. - Pick one offer for v1: Free Audit → $99/mo build & maintain. Subscription is low enough for an ops director's AmEx (no procurement). Per-build pricing ($1.5k–$8k) for users who refuse a recurring relationship. The "Pro Audit" ($499) is a v2 productized SKU for 100–500-person orgs.
- First customer profile: 25–150 person services businesses.Agencies, dev shops, consultancies, fractional ops. Founder/COO buys. Stack is HubSpot/Pipedrive, Slack, Notion/Asana, Gmail, Stripe. Avoid for v1: enterprise (procurement), regulated (PHI/financial advice), <5-person teams (won't pay).
Why this beats the alternatives
- vs. agencies that pitch first, scope later: we lead with value (the audit), so the buyer self-qualifies.
- vs. tools that demand you know what to automate: most operators can't articulate it; the audit does the articulating.
- vs. RPA / Zapier consultancies: we surface the never-automated tier — deals nobody is competing for.
- vs. ChatGPT "audit my business": we have a process database, a scoring rubric, an anti-list, and an opinionated decision tree behind the agent.
The audit job — end-to-end engagement
decay: 6moThe operating spine. Every other section in this doc plugs into one of these six stages.
Six stages
- 1. Intake (60–180s). Company URL, role, size, primary tools. Mode: voice note (default), AI voice call, or typed. Context priming: 8–12 prompts ranging over week-shape, repeated work, hated tasks, last fire-drill.
- 2. Discovery (3–5 min). Conditional branching question bank (§12). Function-aware. Each surfaced pain quantified (frequency × duration × loaded hourly cost) and tagged to a candidate from the process database (§3). Hard caps: max 25 candidates, max 12 minutes total.
- 3. Scoring (~30s, server-side). Each candidate scored on the 7-axis readiness rubric (§13): frequency, time cost, ambiguity, stakes-of-error, data availability, tool-fit, change-mgmt cost. Decision tree (§14) routes each item to: full agent, deterministic workflow, RPA, fix-process-first, or leave-human. Sub-threshold items go to the anti-list explicitly, never silently dropped.
- 4. Recommendation generation (~60s). Top 5 ranked by
dollar_value × readiness × sales_appeal. Each: plain-English explanation, concrete stack, estimated build time, confidence score, similar prior deployments where public. Anti-list section calls out non-recommendations. Headline numbers: total recoverable spend, hours/month. - 5. Deliverable (gated). Email gate unlocks full report. Public preview shows top item only. PDF + shareable URL. Each item has three CTAs: DIY (instructions) · We build for $X · We build + maintain for $99/mo.
- 6. Handoff (sales). Build/subscribe click fires a human-shaped Slack ping with audit + pre-filled proposal. 24-hour SLA. No-pressure path captures email for nurture.
Pricing model — recommendation
Default for v1: Free Audit + $99/mo subscription, with per-build escape hatch ($1.5k–$8k). Subscription is the conversion engine; per-build is the relief valve for buyers who refuse recurring. Comparable: Zapier Experts marketplace per-build rates; emerging "AI ops" subscription shops at $99–$499/mo. zapier.com/experts
Unit economics targets
- Time per audit:<5 min of human time post-MVP (review + send). 100% agent-run is the target. Free audits are only viable if cost-per-audit is durably <$1.50.
- Cost per audit:<$1.50 in LLM/voice/infra at scale. Voice (Retell) ~$0.05/min × 8 min + Claude Sonnet 4.5 ~$0.05 per 50k tokens audit transcript & report = ~$0.65 raw. uncited — modelled, not measured.
- Conversion targets: 4–8% free→$99/mo, 1–3% free→$1.5k+ one-shot.
- LTV target: $1,800 ($99/mo × 18 mo expected SMB tenure).
- CAC target:<$200 via SEO + practitioner presence on X.
Process database
decay: 6moThe 200+ specific, citation-backed automations. Sortable by dollars saved, reliability, sales appeal. Filter by function, status, business size. Default sort favors high-dollar items; flip to status to surface wedges.
Every entry is specific enough that a builder could scope it in one sentence. Status flags: frequently_automated (Zapier/Make territory), rarely_automated (some shops do it, friction exists), never_automated_yet (our wedge tier). Reliability scores carry a ~1.3× haircut relative to vendor claims; see §6 for the rationale. MIT 95%-pilot-failure data
Function
Status
Business size
Min. sales appeal
★ ≥ 1Max. difficulty
★ ≤ 5ofs-020 Scope creep tracking on services contracts Delivery Manager / PM | Ops | mid | 8 | $11,200 | Wedge · never | |||
ofs-017 Shadow IT discovery & cost rationalization IT Procurement / Finance Partner | Ops | mid | 6 | $4,800 | Rarely | |||
chmle-021 Privilege-aware doc summarizer for litigation Litigation Counsel | Legal | enterprise | 40 | $4,800 | Rarely | |||
ofs-002 Vendor renewal alerts before auto-renew Office Manager / IT Ops Lead | Ops | all | 5 | $4,200 | Wedge · never | |||
chmle-022 AI PR review pre-pass (style, security, test coverage) Engineering Manager | Eng | mid | 30 | $3,300 | Frequently | |||
ofs-059 Vendor early-pay discount capture AP Manager / Treasury | Finance | mid | 3 | $3,200 | Wedge · never | |||
chmle-030 Executive inbox triage with action-tagging Executive Assistant / Founder | Exec | all | 18 | $2,700 | Rarely | |||
ofs-097 Voice AI for outbound cold calling SDR / Founder | Sales | all | 60 | $2,520 | Rarely | |||
chmle-027 Internal codebase Q&A bot (Cody-style) Developer Productivity Lead | Eng | mid | 22 | $2,420 | Rarely | |||
chmle-024 On-call alert triage + runbook execution suggestion SRE / DevOps Lead | Eng | mid | 20 | $2,200 | Rarely | |||
chmle-019 Vendor security questionnaire auto-completion Compliance / GRC Lead | Legal | mid | 24 | $2,160 | Rarely | |||
ofs-003 Internal IT ticket triage & first-touch response IT Helpdesk L1 | Ops | mid | 56 | $2,128 | Frequently | |||
chmle-032 Cross-tool exec dashboard (financial + product + people) CEO / CoS | Exec | mid | 14 | $2,100 | Rarely | |||
chmle-020 Regulatory change monitoring + impact briefing Compliance Counsel | Legal | enterprise | 16 | $1,920 | Wedge · never | |||
ofs-051 Board pack assembly (financial section) CFO / FP&A | Finance | mid | 16 | $1,840 | Rarely | |||
ofs-022 Inbound press / partnership / random "interesting" email triage COS / Founder | Ops | all | 16 | $1,760 | Rarely | |||
chmle-029 Weekly board pack auto-draft (metrics + commentary) Chief of Staff | Exec | mid | 16 | $1,760 | Rarely | |||
ofs-084 Account research deep-dive briefs (ABM) AE / ABM Manager | Sales | mid | 24 | $1,728 | Rarely | |||
ofs-042 Revenue recognition rules for SaaS subscriptions Revenue Accountant | Finance | mid | 22 | $1,716 | Rarely | |||
chmle-001 QBR deck auto-generation from product usage + CRM data Customer Success Manager | CS | mid | 24 | $1,680 | Rarely | |||
chmle-005 L1 ticket auto-triage, tag, and KB-grounded draft Support Lead | Support | all | 48 | $1,680 | Frequently | |||
chmle-018 NDA & MSA playbook redlines with attorney HITL Legal Operations Manager | Legal | mid | 18 | $1,620 | Rarely | |||
ofs-092 Demo environment customization per prospect Sales Engineer | Sales | mid | 16 | $1,568 | Rarely | |||
chmle-025 Test flake quarantine + auto-issue Test Infra / Eng | Eng | mid | 14 | $1,540 | Wedge · never | |||
ofs-039 Vendor bill OCR + 3-way match AP Specialist | Finance | mid | 42 | $1,512 | Frequently | |||
ofs-075 Outbound email personalization at scale SDR / Founder | Sales | all | 36 | $1,512 | Frequently | |||
ofs-068 AP fraud / duplicate invoice detection AP / Controller | Finance | mid | 3 | $1,500 | Rarely | |||
chmle-031 Calendar prep briefs (who, why, last touch) EA / Founder | Exec | all | 10 | $1,500 | Rarely | |||
ofs-026 Vendor risk reassessment annual cycle Vendor Risk Manager / Compliance | Ops | enterprise | 18 | $1,476 | Rarely | |||
vrt-040 Proactive shipment status comms Operations / customer service rep (CSR) · logistics | Vertical | all | 40 | $1,400 | Frequently | |||
ofs-050 Audit prep PBC (provided-by-client) list management Controller | Finance | mid | 18 | $1,296 | Rarely | |||
chmle-002 Customer health score rollup with leading-indicator churn flags CS Ops Lead | CS | mid | 18 | $1,260 | Wedge · never | |||
chmle-015 Ad creative iteration from winning concepts Performance Marketer | Marketing | SMB | 18 | $1,260 | Rarely | |||
ofs-007 Contract redline coordination & version tracking Legal Ops / COO | Ops | mid | 18 | $1,224 | Rarely | |||
ofs-041 Monthly close checklist orchestration Controller / Accounting Manager | Finance | mid | 18 | $1,224 | Rarely | |||
ofs-010 Cross-team status update aggregation PMO / Chief of Staff | Ops | mid | 20 | $1,160 | Rarely | |||
ofs-018 Customer health score rollup from product telemetry + support tickets RevOps / CS Ops | Ops | mid | 16 | $1,152 | Wedge · never | |||
ofs-016 Compliance evidence collection for SOC2 / ISO Security / Compliance Lead | Ops | mid | 14 | $1,148 | Frequently | |||
vrt-050 Document request list (DRL) chase Senior associate / coordinator · professional_services | Vertical | all | 20 | $1,100 | Frequently | |||
chmle-012 Internal HR Q&A bot (benefits, policy, PTO) HR Operations | HR | mid | 22 | $1,100 | Rarely | |||
ofs-044 Cash flow forecasting (13-week rolling) FP&A / Controller | Finance | mid | 15 | $1,080 | Rarely | |||
vrt-025 Buyer showing scheduling across multiple listings Agent / ISA · real_estate | Vertical | SMB | 12 | $1,080 | Rarely | |||
vrt-049 Conflict check on new matter Conflicts attorney / partner · professional_services | Vertical | all | 8 | $1,080 | Frequently | |||
vrt-032 Prior auth submission and status follow-up PA coordinator · healthcare | Vertical | all | 30 | $1,050 | Rarely | |||
vrt-033 Denial triage and resubmission Biller / RCM lead · healthcare | Vertical | all | 30 | $1,050 | Rarely | |||
vrt-044 Driver check-in calls / status capture Dispatcher · logistics | Vertical | SMB | 30 | $1,050 | Rarely | |||
chmle-004 Renewal forecasting 90/60/30 day playbook Customer Success Director | CS | mid | 12 | $1,050 | Rarely | |||
ofs-077 Meeting prep brief generation AE / Manager | Sales | all | 18 | $1,044 | Rarely | |||
ofs-004 Employee onboarding asset provisioning IT Ops / People Ops | Ops | mid | 24 | $1,008 | Frequently | |||
vrt-003 Statement of Work generation from sales call Founder / new biz lead · agencies | Vertical | SMB | 10 | $1,000 | Rarely | |||
vrt-010 First-pass RFP response from past wins New biz / founder · ecommerce | Vertical | all | 10 | $1,000 | Wedge · never | |||
chmle-008 Recruiter outreach personalization at scale Recruiter | Recruiting | mid | 20 | $1,000 | Frequently | |||
chmle-010 Onboarding day-1-to-30 task orchestrator People Ops Manager | HR | all | 16 | $992 | Rarely | |||
ofs-060 Intercompany reconciliation (multi-entity) Controller / GL Accountant | Finance | enterprise | 12 | $984 | Rarely | |||
ofs-079 Deal coaching from call transcripts Sales Manager | Sales | mid | 12 | $984 | Rarely | |||
chmle-009 Interview scheduling across panel + candidate availability Recruiting Coordinator | Recruiting | mid | 28 | $980 | Frequently | |||
chmle-014 Lifecycle email triggers from product behavior Lifecycle Marketing Manager | Marketing | mid | 14 | $980 | Frequently | |||
ofs-089 CRM data hygiene (missing fields, bad data) RevOps | Sales | mid | 14 | $952 | Rarely | |||
ofs-001 Vendor onboarding packet collection Procurement / Vendor Manager | Ops | mid | 22 | $924 | Rarely | |||
vrt-052 Monthly close + advisory package Senior accountant · professional_services | Vertical | SMB | 10 | $900 | Rarely | |||
ofs-045 Budget vs. actual variance flagging FP&A Analyst | Finance | mid | 13 | $884 | Rarely | |||
vrt-005 KPI dashboard build/refresh for client portal Analyst · agencies | Vertical | all | 16 | $880 | Frequently | |||
chmle-013 SEO content brief generator from SERP + competitor scrape Content Strategist | Marketing | SMB | 16 | $880 | Frequently | |||
chmle-017 Repurposing long-form video → 10 short-form clips Content Producer | Marketing | SMB | 22 | $880 | Frequently | |||
vrt-011 Where is my order" support ticket resolution Support lead · ecommerce | Vertical | SMB | 25 | $875 | Frequently | |||
vrt-027 Tenant maintenance request triage and dispatch Maintenance coordinator · real_estate | Vertical | all | 25 | $875 | Rarely | |||
vrt-031 Pre-visit insurance + benefits verification Front desk / billing · healthcare | Vertical | all | 25 | $875 | Rarely | |||
vrt-043 Customer rate quote response Sales / pricing analyst · logistics | Vertical | SMB | 25 | $875 | Rarely | |||
ofs-009 Weekly all-hands KPI rollup COS / Operations Manager | Ops | all | 14 | $868 | Rarely | |||
ofs-076 Outbound follow-up sequencing SDR / AE | Sales | all | 18 | $864 | Frequently | |||
ofs-081 Churn risk flagging for sales-led GTM CSM / AE | Sales | mid | 12 | $864 | Rarely | |||
ofs-103 Slack channel for shared customer (Slack Connect) AE / CSM | Sales | mid | 12 | $864 | Wedge · never | |||
chmle-007 Voice-of-customer rollup from tickets + surveys + reviews CX Lead | Support | mid | 12 | $840 | Rarely | |||
ofs-011 Exception monitoring for ops dashboards Operations Analyst | Ops | mid | 16 | $832 | Rarely | |||
ofs-080 Pipeline hygiene audit RevOps / Sales Manager | Sales | mid | 10 | $820 | Rarely | |||
ofs-014 BD / partnership pipeline reporting Head of Partnerships / BD Ops | Ops | mid | 12 | $816 | Wedge · never | |||
ofs-052 FP&A model refresh from source systems FP&A Analyst | Finance | mid | 12 | $816 | Rarely | |||
ofs-008 Internal wiki staleness detection COS / Knowledge Manager | Ops | all | 14 | $812 | Wedge · never | |||
ofs-078 Post-call CRM update from recording AE | Sales | all | 14 | $812 | Frequently | |||
ofs-038 AR follow-up cadence on overdue invoices AR Specialist / Controller | Finance | all | 19 | $798 | Rarely | |||
ofs-012 Capacity planning for services teams Resource Manager / Delivery Lead | Ops | mid | 11 | $792 | Wedge · never | |||
ofs-073 Lead enrichment from form fill SDR / RevOps | Sales | all | 18 | $756 | Frequently | |||
ofs-098 Inbound voice qualification (chatbot/voicebot) SDR / Marketing | Sales | all | 18 | $756 | Rarely | |||
ofs-082 Quote / proposal generation AE / Sales Engineer | Sales | all | 12 | $744 | Rarely | |||
ofs-083 Contract redline negotiation responses AE / Legal / RevOps | Sales | mid | 9 | $738 | Rarely | |||
ofs-090 Sales tax / pricing approval workflow Deal Desk / RevOps | Sales | mid | 10 | $720 | Rarely | |||
vrt-022 CMA report generation for seller meeting Agent · real_estate | Vertical | SMB | 8 | $720 | Rarely | |||
vrt-051 Daily time entry reconstruction Associate / consultant · professional_services | Vertical | all | 8 | $720 | Rarely | |||
chmle-023 PR description + changelog generation from diff Engineering IC | Eng | all | 8 | $720 | Frequently | |||
chmle-026 Dependency update PRs with safety gates Engineering | Eng | all | 8 | $720 | Rarely | |||
vrt-036 Digital pre-visit intake (history, insurance, consents) Front desk · healthcare | Vertical | SMB | 20 | $700 | Frequently | |||
vrt-041 Proof of delivery collection from carriers Settlement clerk · logistics | Vertical | SMB | 20 | $700 | Rarely | |||
chmle-006 KB article auto-generation from resolved tickets Support Ops | Support | mid | 14 | $700 | Wedge · never | |||
ofs-067 Treasury — daily cash position consolidation Treasurer / Controller | Finance | mid | 9 | $675 | Rarely | |||
ofs-040 Expense receipt categorization from photos Controller / AP | Finance | all | 16 | $672 | Frequently | |||
vrt-001 Weekly client status report assembly Account manager · agencies | Vertical | all | 12 | $660 | Frequently | |||
vrt-002 Time entry compliance and reconstruction Ops manager / agency owner · agencies | Vertical | all | 12 | $660 | Rarely | |||
vrt-004 Monthly social/blog content calendar with hooks + topics Content strategist · agencies | Vertical | SMB | 12 | $660 | Rarely | |||
vrt-014 Meta/TikTok creative testing rotation + winner-killer logic Media buyer · ecommerce | Vertical | mid | 12 | $660 | Rarely | |||
vrt-016 Inventory replenishment alerting Ops manager · ecommerce | Vertical | SMB | 12 | $660 | Frequently | |||
chmle-028 Incident postmortem first draft Engineering Manager | Eng | mid | 6 | $660 | Rarely | |||
ofs-101 Renewal & expansion forecast for sales-assisted CS CSM / Account Manager | Sales | mid | 9 | $648 | Rarely | |||
ofs-005 Employee offboarding access revocation audit IT Security / People Ops | Ops | all | 12 | $624 | Rarely | |||
ofs-019 Internal data hygiene drift detection RevOps / Data Analyst | Ops | mid | 10 | $620 | Wedge · never | |||
ofs-062 Subscription billing reconciliation (CRM vs. Stripe) RevOps / Controller | Finance | all | 10 | $620 | Wedge · never | |||
chmle-016 Brand mention monitoring with sentiment + auto-response queue Social Media Manager | Marketing | SMB | 12 | $600 | Frequently | |||
ofs-088 Reply triage from outbound sequences SDR / AE | Sales | all | 14 | $588 | Frequently | |||
ofs-086 Win/loss analysis from CRM + call data RevOps / Product Marketing | Sales | mid | 8 | $576 | Wedge · never | |||
vrt-020 Cold outreach to UGC creators + deliverable tracking Influencer manager · real_estate | Vertical | all | 16 | $560 | Rarely | |||
vrt-037 Outbound referral tracking + records request Referral coordinator · healthcare | Vertical | SMB | 16 | $560 | Wedge · never | |||
chmle-003 Expansion-signal detection from feature adoption + seat growth Account Manager | CS | mid | 8 | $560 | Wedge · never | |||
ofs-070 Procurement card reconciliation across multiple GLs AP / Controller | Finance | mid | 13 | $546 | Frequently | |||
ofs-074 Outbound list building from ICP SDR / Growth | Sales | all | 13 | $546 | Frequently | |||
vrt-021 MLS listing description + headline Listing agent / coordinator · real_estate | Vertical | SMB | 6 | $540 | Rarely | |||
vrt-030 Fair Housing + MLS compliance check on draft listings Compliance / managing broker · healthcare | Vertical | SMB | 6 | $540 | Wedge · never | |||
vrt-053 Filing deadline calendar for regulated clients Compliance / paralegal · professional_services | Vertical | all | 6 | $540 | Wedge · never | |||
vrt-054 Firm-wide precedent / past-work retrieval All knowledge workers · professional_services | Vertical | all | 6 | $540 | Rarely | |||
vrt-012 Return request triage and authorization Support manager · ecommerce | Vertical | SMB | 15 | $525 | Frequently | |||
vrt-013 Long-form product description from PDP brief Merchant/content · ecommerce | Vertical | SMB | 15 | $525 | Frequently | |||
vrt-026 Buyer/seller closing document chase Transaction coordinator (TC) · real_estate | Vertical | SMB | 15 | $525 | Rarely | |||
vrt-034 No-show prevention + same-day waitlist fill Scheduler · healthcare | Vertical | SMB | 15 | $525 | Frequently | |||
vrt-042 Post-delivery invoice audit vs quote AP / freight auditor · logistics | Vertical | all | 15 | $525 | Frequently | |||
vrt-045 Customer-facing ETA on last-mile delivery Dispatch / CS · logistics | Vertical | all | 15 | $525 | Frequently | |||
vrt-057 Daily route optimization for technicians Dispatcher · local_services | Vertical | SMB | 15 | $525 | Frequently | |||
vrt-064 Daily crew confirmation + replacement when out Operations · local_services | Vertical | SMB | 15 | $525 | Wedge · never | |||
ofs-025 New client kickoff packet generation Project Manager / CS Lead | Ops | all | 9 | $522 | Rarely | |||
ofs-032 Internal expense policy enforcement Controller / People Ops | Ops | mid | 8 | $496 | Rarely | |||
ofs-034 Recurring-meeting agenda hygiene COS / Manager | Ops | mid | 8 | $496 | Wedge · never | |||
ofs-023 Time tracking enforcement for billable teams Delivery Ops / Controller | Ops | mid | 10 | $480 | Rarely | |||
ofs-069 SaaS metrics calculation (NRR, GRR, LTV, CAC) FP&A / RevOps | Finance | mid | 7 | $476 | Frequently | |||
ofs-094 Partner referral attribution Partnerships / RevOps | Sales | mid | 7 | $476 | Wedge · never | |||
ofs-096 NPS / CSAT response triage to sales actions CSM / RevOps | Sales | mid | 8 | $464 | Wedge · never | |||
ofs-053 Equity / cap table waterfall scenario modeling CFO / Founder | Finance | mid | 4 | $460 | Rarely | |||
vrt-024 Long-term nurture for warm but not ready buyers Agent · real_estate | Vertical | SMB | 5 | $450 | Frequently | |||
vrt-048 New client engagement letter Partner / senior · professional_services | Vertical | SMB | 5 | $450 | Frequently | |||
vrt-007 Real-time retainer hour tracking + scope-creep flagging Project manager · agencies | Vertical | all | 8 | $440 | Rarely | |||
ofs-085 Mutual action plan (MAP) generation & tracking AE | Sales | mid | 7 | $434 | Wedge · never | |||
ofs-043 Weekly Stripe revenue rollup into Notion + Slack Founder / Controller | Finance | SMB | 5 | $425 | Rarely | |||
ofs-054 Bank reconciliation Bookkeeper / Controller | Finance | all | 10 | $420 | Frequently | |||
vrt-006 Creative asset versioning and handoff to media buyer Producer · agencies | Vertical | all | 12 | $420 | Rarely | |||
vrt-023 Buyer/seller lead routing + first-touch Team lead · real_estate | Vertical | all | 12 | $420 | Frequently | |||
vrt-029 Post vacancy to 20+ rental sites Leasing · real_estate | Vertical | SMB | 12 | $420 | Frequently | |||
vrt-038 Patient-responsibility AR follow-up Biller · healthcare | Vertical | SMB | 12 | $420 | Rarely | |||
vrt-055 Vendor invoice approval and posting Controller / bookkeeper · local_services | Vertical | all | 12 | $420 | Frequently | |||
vrt-056 Inbound lead capture + first response Owner / dispatcher · local_services | Vertical | SMB | 12 | $420 | Rarely | |||
chmle-011 Performance review prep aggregator (PRs, 1:1s, peer notes) HRBP | HR | mid | 6 | $420 | Rarely | |||
ofs-013 Meeting scheduling across multiple parties EA / Office Manager | Ops | all | 13 | $416 | Frequently | |||
ofs-058 R&D tax credit documentation collection Controller / Tax Advisor | Finance | all | 5 | $410 | Rarely | |||
ofs-029 Service-level SLA breach detection Customer Success / Support Ops | Ops | mid | 7 | $406 | Rarely | |||
ofs-105 Customer reference matching for sales calls AE / Customer Marketing | Sales | mid | 7 | $406 | Rarely | |||
ofs-006 Laptop & monitor asset tracking IT Ops / Office Manager | Ops | mid | 9 | $378 | Rarely | |||
ofs-064 Employee corporate card spend review Controller / Manager | Finance | all | 6 | $372 | Rarely | |||
ofs-095 Multi-thread champion tracking AE / RevOps | Sales | mid | 6 | $372 | Wedge · never | |||
ofs-036 Procurement spend categorization Procurement Analyst | Ops | mid | 7 | $364 | Rarely | |||
ofs-063 Customer prepayment / deferred revenue tracking Revenue Accountant | Finance | all | 5 | $360 | Rarely | |||
vrt-028 Rent reminder/late notice cadence Property manager · real_estate | Vertical | SMB | 10 | $350 | Frequently | |||
vrt-035 Patient recall / recare outreach Front desk · healthcare | Vertical | SMB | 10 | $350 | Frequently | |||
vrt-039 Provider credentialing renewals + payer enrollment Credentialing coordinator · logistics | Vertical | all | 10 | $350 | Wedge · never | |||
vrt-047 New carrier onboarding for broker Carrier rep · professional_services | Vertical | SMB | 10 | $350 | Frequently | |||
vrt-059 On-site quote and invoice generation Technician · local_services | Vertical | SMB | 10 | $350 | Rarely | |||
vrt-063 Permit-ready job photo + doc packet Project coordinator · local_services | Vertical | SMB | 10 | $350 | Wedge · never | |||
ofs-028 Internal NPS / pulse survey loop People Ops | Ops | mid | 6 | $348 | Rarely | |||
ofs-031 NDA / MSA template generation for new prospects Sales Ops / Legal | Ops | all | 6 | $348 | Frequently | |||
ofs-037 Internal reference check for hiring Recruiter / Hiring Manager | Ops | mid | 6 | $348 | Wedge · never | |||
ofs-047 Sales tax nexus monitoring & filing prep Controller / Accountant | Finance | all | 6 | $348 | Frequently | |||
ofs-066 Customer pricing tier audit & uplift opportunity RevOps / Finance Partner | Finance | mid | 5 | $340 | Wedge · never | |||
vrt-008 New-project kickoff brief from CRM + discovery PM · agencies | Vertical | SMB | 6 | $330 | Rarely | |||
ofs-057 Equity refresh grants & vesting acceleration tracking People Ops / CFO | Finance | mid | 4 | $328 | Wedge · never | |||
ofs-055 Foreign exchange revaluation Controller / Treasury | Finance | mid | 4 | $312 | Rarely | |||
ofs-065 Statement-of-cash-flows (indirect method) prep Controller / FP&A | Finance | mid | 4 | $312 | Rarely | |||
ofs-100 Competitor mention alerts from review sites Product Marketing / CI | Sales | mid | 5 | $310 | Rarely | |||
ofs-102 Discovery call agenda customization AE | Sales | all | 5 | $310 | Wedge · never | |||
ofs-107 Champion enablement asset delivery AE / Customer Marketing | Sales | mid | 5 | $310 | Wedge · never | |||
ofs-056 Customer payment failure recovery AR / RevOps | Finance | all | 7 | $294 | Frequently | |||
ofs-027 Conference / event ROI tracking Field Marketing / Ops | Ops | mid | 5 | $290 | Wedge · never | |||
ofs-046 Payroll variance check vs. prior period Payroll Manager / Controller | Finance | mid | 5 | $290 | Wedge · never | |||
ofs-087 ICP scoring on inbound leads RevOps / Marketing Ops | Sales | mid | 5 | $290 | Rarely | |||
ofs-071 Customer-by-customer profitability analysis FP&A | Finance | mid | 4 | $288 | Wedge · never | |||
vrt-009 Hours-to-invoice reconciliation Bookkeeper/ops · agencies | Vertical | SMB | 8 | $280 | Rarely | |||
vrt-018 Per-customer post-purchase flow personalization Lifecycle / CRM marketer · ecommerce | Vertical | all | 8 | $280 | Rarely | |||
vrt-046 Damage / loss claim filing with carriers Claims clerk · logistics | Vertical | SMB | 8 | $280 | Wedge · never | |||
vrt-058 Technician ETA + on-the-way SMS Tech / dispatcher · local_services | Vertical | SMB | 8 | $280 | Frequently | |||
vrt-061 Maintenance contract renewal + seasonal reminders Office manager · local_services | Vertical | SMB | 8 | $280 | Frequently | |||
ofs-099 SDR-to-AE handoff brief SDR / AE | Sales | mid | 5 | $260 | Wedge · never | |||
ofs-049 Customer credit memo issuance AR / Controller | Finance | all | 6 | $252 | Rarely | |||
ofs-061 Equity-based comp accrual (ASC 718) Revenue/Equity Accountant | Finance | mid | 3 | $246 | Rarely | |||
ofs-072 Lease accounting (ASC 842) Controller / Lease Accountant | Finance | mid | 3 | $246 | Frequently | |||
ofs-091 Trigger-based outreach (job change, funding, hiring) SDR / Growth | Sales | all | 5 | $240 | Rarely | |||
ofs-033 Holiday / leave coverage planning Manager / People Ops | Ops | mid | 4 | $232 | Wedge · never | |||
ofs-024 Real estate / office maintenance ticket coordination Facilities / Office Manager | Ops | mid | 6 | $216 | Rarely | |||
ofs-108 Lost-deal interview scheduling & analysis Product Marketing | Sales | mid | 3 | $216 | Wedge · never | |||
vrt-015 Review response across review sites Brand/CX manager · ecommerce | Vertical | SMB | 6 | $210 | Frequently | |||
vrt-017 Fraud review on flagged orders before fulfillment Ops manager · ecommerce | Vertical | all | 6 | $210 | Frequently | |||
vrt-062 Parts inventory reorder Warehouse / ops · local_services | Vertical | SMB | 6 | $210 | Rarely | |||
vrt-065 COI + W9 collection from subs Office manager · local_services | Vertical | SMB | 6 | $210 | Rarely | |||
vrt-066 We sent you a quote 5 days ago" follow-up Owner / sales | Vertical | SMB | 6 | $210 | Rarely | |||
ofs-021 PTO conflict detection People Ops / Manager | Ops | all | 4 | $208 | Wedge · never | |||
ofs-015 Office supplies & snack reorder Office Manager | Ops | all | 6 | $192 | Frequently | |||
ofs-093 Lost deal re-engagement campaigns Marketing / RevOps | Sales | mid | 3 | $174 | Wedge · never | |||
ofs-104 Pricing change communication to existing customers AE / CS / Marketing | Sales | all | 3 | $174 | Wedge · never | |||
ofs-030 Insurance certificate (COI) collection from vendors Risk / Procurement | Ops | mid | 4 | $168 | Wedge · never | |||
ofs-048 1099 contractor compliance & filing Controller / AP | Finance | all | 3 | $145 | Frequently | |||
ofs-106 ICP refinement from closed-won analysis RevOps / Product Marketing | Sales | mid | 2 | $144 | Wedge · never | |||
vrt-019 Quarterly Shopify app audit + consolidation COO / ops · ecommerce | Vertical | SMB | 4 | $140 | Wedge · never | |||
vrt-060 Post-job review request Owner · local_services | Vertical | SMB | 4 | $140 | Frequently | |||
ofs-035 Domain renewal & DNS monitoring IT Ops / Marketing Ops | Ops | all | 1 | $58 | Rarely |
Wedge — top processes & never_automated tier
decay: 6moThe 10 deals to lead with, plus the wedge tier of pains nobody else is solving yet.
Top 10 highest-ROI starting processes
Ranked by cost_saved × sales_appeal × reliability_today. These are the deals that should populate cold-outbound proof points and landing page case studies first.
| # | Process | Function | $/mo | Rel | Sell |
|---|---|---|---|---|---|
| 01 | Scope creep tracking on services contracts | Ops | $11,200 | 3 | 5 |
| 02 | Shadow IT discovery & cost rationalization | Ops | $4,800 | 4 | 5 |
| 03 | Vendor renewal alerts before auto-renew | Ops | $4,200 | 4 | 5 |
| 04 | AI PR review pre-pass (style, security, test coverage) | Eng | $3,300 | 4 | 5 |
| 05 | Privilege-aware doc summarizer for litigation | Legal | $4,800 | 3 | 4 |
| 06 | Executive inbox triage with action-tagging | Exec | $2,700 | 4 | 5 |
| 07 | Vendor early-pay discount capture | Finance | $3,200 | 4 | 4 |
| 08 | Vendor security questionnaire auto-completion | Legal | $2,160 | 4 | 5 |
| 09 | Cross-tool exec dashboard (financial + product + people) | Exec | $2,100 | 4 | 5 |
| 10 | Internal codebase Q&A bot (Cody-style) | Eng | $2,420 | 4 | 4 |
The wedge tier — never_automated_yet
These are the processes operators complain about, but no vendor or agency has packaged a clean solution for. Low competitive density. Cheap to acquire customers. Hard for a competitor to copy without our process database.
- Vendor renewal alerts before auto-renew — Never get blindsided by a $40k auto-renewal again.
- Internal wiki staleness detection — Stop pointing new hires at lies.
- Capacity planning for services teams — Know in week 2 you'll be underwater in week 6, not week 5.
- BD / partnership pipeline reporting — Partnerships are an afterthought because reporting is. Fix the reporting.
- Customer health score rollup from product telemetry + support tickets — Know which account will churn this quarter — before the cancel email.
- Internal data hygiene drift detection — Your dashboard isn't lying to you anymore.
- Scope creep tracking on services contracts — Stop eating $20k of scope creep per project.
- PTO conflict detection — Don't be the manager who let 3 PMs go to Italy the same week.
- Conference / event ROI tracking — Know which conferences actually drove revenue, not just leads.
- Insurance certificate (COI) collection from vendors — COIs stop being a fire drill before a big event.
- Holiday / leave coverage planning — December coverage planning is 10 minutes, not a panic.
- Recurring-meeting agenda hygiene — Meetings get useful again.
Anti-list — don't automate this
decay: 6mo50+ processes that look automatable but bite. Filter by failure mode. Selling the anti-list is a credibility flex no vendor offers.
Failure modes draw on real production incidents: Air Canada chatbot making up refund policy BBC ; Klarna walking back AI-first support Customer Experience Dive ; NYC Local Law 144 on automated employment decision tools NYC DCWP ; CVE-2025-59944 prompt-injection-to-RCE in Cursor mayhemcode; OpenAI dropping SWE-bench Verified for benchmark contamination morphllm.
60 of 60 traps
Fully automated customer refunds / chatbot refund policy
Why it tempts you
Refunds are repetitive and seem rule-based
Concrete failures
regulation, judgment Air Canada chatbot invented a bereavement-fare refund window; BC Civil Resolution Tribunal held Air Canada liable for negligent misrepresentation by AI; $650 CAD damages (https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416, https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/, Feb 2024). Tribunal explicitly rejected "chatbot is a separate entity" defense. Cursor support bot fabricated a new "one-device login" policy in April 2025; customers cancelled subscriptions; company refunded and apologized publicly (https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/, https://fortune.com/article/customer-support-ai-cursor-went-rogue/).
Right split
AI drafts response + cites policy URL; agent (or rule engine on narrow, well-bounded refund types like "unworn item < 30 days") approves payment. Refund authority above $X never automatic.
Full email reply automation without HITL
Why it tempts you
Inbox is the timesuck owners complain about most
Concrete failures
judgment, brand-voice, hallucination, relationship-dependent DPD chatbot, after Jan 2024 update, swore at customer, wrote a poem mocking DPD, called itself the worst delivery firm in the world; 1.3M views before disabled (https://time.com/6564726/ai-chatbot-dpd-curses-criticizes-company/). Klarna replaced 700 CS agents with AI, then in 2025 admitted "we went too far," now rehiring; CEO public statement (https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396, https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/).
Right split
AI drafts in agent inbox; agent edits 30s and sends. Auto-send only for `intent=order_status` lookups with order data validated.
AI cold sales calls from outbound dialers
Why it tempts you
Looks like 100x SDR leverage
Concrete failures
regulation, brand FCC Declaratory Ruling Feb 8, 2024 confirms TCPA's "artificial or prerecorded voice" prohibition includes AI-generated voices; prior express consent required (https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal). Civil penalties up to $1,500 per call per recipient. Class action exposure under TCPA is severe — single campaign can hit eight-figure damages.
Right split
AI for inbound voice qualification, opted-in re-engagement, internal calls (recruiter to candidate after consent). Cold outbound stays human or written channels.
AI moderation at scale without escalation paths
Why it tempts you
Massive volume, repetitive
Concrete failures
judgment, low-frequency-high-stakes, regulation DSA in EU mandates human review options for content decisions affecting users; pure AI moderation breaches Art 14/20 process rights. Repeated Meta/YouTube false-positives banning veterans accounts, breast cancer support groups, etc. — recurring brand and PR risk.
Right split
AI as first triage and high-confidence auto-action on clear violations (CSAM, spam); human reviews any escalations and user appeals within 24h.
Stripe dispute responses fully automated
Why it tempts you
Dispute evidence packets are structured
Concrete failures
judgment, regulation, cost Stripe Smart Disputes (their own AI tool) takes 30% of recovered amount and only beats manual on small disputes (https://directpaynet.com/stripe-forcing-ai-dispute-tool-taking-30-of-winnings/, 2025). Manual dispute response has <20% win rate; bad automated submissions can lock you out of resubmitting (https://www.chargeflow.io/blog/stripe-dispute-fees-2025). Fraud signaling that turns out to be "friendly fraud" needs nuance no LLM has yet.
Right split
AI assembles evidence packet + draft narrative; ops manager reviews for high-value disputes; auto-submit only for low-value, low-complexity friendly fraud where evidence is air-tight.
AI press release / crisis comms drafting (auto-publish)
Why it tempts you
Templated, low-frequency, looks safe
Concrete failures
low-frequency-high-stakes, judgment, brand One factual error in a crisis statement compounds 100x; cleanup cost > entire annual comms budget. LLMs frequently hallucinate exec quotes and unverified incident details.
Right split
AI drafts FAQ + holding statement template against a `crisis_playbook.md`; PR lead and legal sign before publish, always.
AI hiring decisions (resume screening + auto-reject)
Why it tempts you
1,000 resumes per req, screening is a slog
Concrete failures
regulation (Title VII, ADEA, ADA, NYC LL144, EU AI Act, Illinois AI Video Interview Act, Colorado AI Act), judgment Mobley v. Workday: certified as nationwide ADEA collective action May 2025; Workday admitted 1.1B applications rejected by tool in scope period (https://www.lawandtheworkplace.com/2025/06/ai-bias-lawsuit-against-workday-reaches-next-stage-as-court-grants-conditional-certification-of-adea-claim/, https://www.fisherphillips.com/en/insights/insights/discrimination-lawsuit-over-workdays-ai-hiring-tools-can-proceed-as-class-action-6-things). iTutorGroup paid $365K to settle EEOC's first AI-discrimination case; software auto-rejected applicants over 55/60 based on birth date (https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit, Aug 2023). NYC Local Law 144 since July 2023 requires bias audits + candidate notice; $500-$1,500/violation/day (https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page). EU AI Act classifies HR AI as high-risk; full requirements enforceable Aug 2, 2026; fines up to €35M or 7% global turnover (https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai).
Right split
AI extracts and standardizes resume data + ranks by stated criteria; humans make every interview/reject decision. Audit logs of every AI input/output. Notice to candidates per NYC LL144 + EU AI Act.
AI performance reviews
Why it tempts you
Managers procrastinate on reviews; AI seems "objective
Concrete failures
judgment, relationship-dependent, regulation EU AI Act high-risk: "AI systems used to evaluate workers' performance" explicitly covered (Annex III) -> mandatory risk assessment, human oversight, bias testing. California AB 2930 and Colorado AI Act both name performance-management as covered. Empirical: AI summaries based on Slack/email volume penalize quieter contributors, ESL workers, and parents on flex schedules — disparate impact risk.
Right split
AI aggregates ticket counts, code review feedback, peer comments into a tab; manager owns the narrative and decision. Never auto-rate.
AI for sensitive HR conversations (PIPs, terminations)
Why it tempts you
Managers want a script
Concrete failures
regulation, judgment, relationship-dependent WARN Act notices need legal-grade precision; Title VII / ADA / FMLA traps in seemingly neutral language. Recorded AI conversations with employees raise wiretap concerns (two-party consent states).
Right split
AI prepares HR talking-point checklist for manager pre-meeting; conversations are always human. AI never present in the room or scripting in real time.
AI lead/talent screening without feedback loop
Why it tempts you
Score everyone, focus reps on top
Concrete failures
data-quality, regulation Model drift after 6 months on stale CRM labels; reps stop trusting it, ROI vanishes. Reverse causation: high-scored leads get more rep effort -> appear to convert better -> model self-validates a false signal.
Right split
AI suggests; rep accepts/rejects with reason; reason-codes feed weekly retrain. Never auto-suppress leads from human visibility.
AI legal redlines without attorney review
Why it tempts you
NDA review is repetitive
Concrete failures
regulation (UPL), low-frequency-high-stakes, judgment Damien Charlotin's hallucination database tracks 1,348+ documented cases of AI-fabricated citations in court filings by mid-2026 (https://www.damiencharlotin.com/hallucinations/). Stanford RegLab found Lexis+ AI hallucinated >17%, Westlaw AI-Assisted Research >34% of the time (https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries). ABA Formal Opinion 512 (July 29, 2024) imposes verification duty on lawyers.
Right split
AI flags clauses against firm playbook + drafts proposed redlines; attorney must read and sign every edit. Never send out a redline the bot wrote untouched.
Auto-renewals of vendor / customer contracts
Why it tempts you
Reduces churn / spend creep
Concrete failures
change-management, judgment, regulation (state auto-renewal laws — CA SB-313, NY GBL 5-903) California, NY, Illinois, and others require explicit re-disclosure before auto-renewal; FTC Click-to-Cancel rule pending. Auto-renewing without proper notice is grounds for restitution. Vendor cost creep: SaaS bill grows 15-20%/yr if nobody pushes back.
Right split
AI surfaces upcoming renewals 60/30/14 days out with usage data and proposed action; ops decides renew/renegotiate/cancel.
AI auto-translation of legal / medical content
Why it tempts you
Translation is "easy
Concrete failures
regulation, low-frequency-high-stakes, judgment HHS Section 1557 (ACA) requires qualified medical interpreters; FDA labeling errors trigger recalls. EU MDR / IVDR require accuracy attestations on translated medical device IFUs.
Right split
AI first draft + glossary enforcement; certified linguist signs off. Never auto-publish in regulated domains.
Generated marketing claims without legal review (FTC)
Why it tempts you
Scale ad creative with LLMs
Concrete failures
regulation FTC's Operation AI Comply (launched Sept 2024) has brought 12+ enforcement actions in 2025 against AI-washing and unsubstantiated AI-product claims (https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes). DoNotPay fined $193K for unsubstantiated AI claims; Workado settled for "98% accuracy" claim it could not substantiate.
Right split
AI drafts ad copy with no superlative/health/earnings claims allowed; legal reviews any claim invoking safety, efficacy, or earnings.
AI medical advice / clinical decision support without FDA review
Why it tempts you
Tempting to use LLM in patient-facing chat
Concrete failures
regulation (FDA SaMD), low-frequency-high-stakes, privacy FDA's December 2024 draft guidance on AI-enabled device software functions; "clinical decision support" function may require 510(k). HIPAA + state telehealth laws bar unlicensed practice.
Right split
AI scheduling, intake, FAQ, post-visit summaries — no diagnosis, no medication advice, no triage. Always escalate clinical questions to clinician.
Bookkeeping reconciliation fully automated (no CPA gate)
Why it tempts you
Bank feeds + GPT can categorize most transactions
Concrete failures
regulation, data-quality, judgment Mis-categorizations compound; year-end tax cleanup costs 3-5x what catching them monthly would have. Sales tax nexus questions need CPA judgment.
Right split
AI categorizes routine transactions (matching prior patterns) + flags uncertain ones; bookkeeper reviews weekly; CPA signs financials.
Tax preparation fully automated for non-trivial returns
Why it tempts you
1040EZ is solved; why not Schedule C?
Concrete failures
regulation (Circular 230), low-frequency-high-stakes IRS Circular 230 imposes preparer due diligence; Penalty $635/return for unreasonable positions. State-specific items (e.g., NY decoupling from federal bonus depreciation) trip LLMs.
Right split
AI pre-populates Schedule entries from intake docs; CPA reviews and signs. Never auto-file.
Full data deletion / GDPR right-to-erasure agent
Why it tempts you
DSAR volume is annoying and growing
Concrete failures
regulation, auditability, vendor-lock GDPR Art 17 and CCPA both require evidence-of-completion; California Delete Act mandates triennial independent audit from Jan 2028 (https://www.didomi.io/blog/california-delete-act). Agents that "delete" but skip backup tiers, archived emails, ML training datasets — leave non-compliant residual data.
Right split
AI orchestrates deletion workflow across systems and produces evidence pack; privacy officer reviews and signs deletion certificate before reply to data subject.
Crisis comms auto-draft and send
Why it tempts you
Speed matters in crises
Concrete failures
low-frequency-high-stakes, judgment Inaccurate "facts" in initial statement become the story (e.g., wrong casualty figures, wrong fault attribution).
Right split
AI prepares holding-statement template + journalist Q&A docs based on factual brief; CEO/PR/legal approve every word that exits.
Pricing changes via agent
Why it tempts you
Dynamic pricing seems like easy margin
Concrete failures
cost, judgment, regulation Wendy's "surge pricing" 2024 PR backlash forced retraction within days. Amazon's 2011 fly-genetics book priced at $23M by competing bots (no humans noticing). Some EU member states bar personalized pricing without disclosure.
Right split
AI surfaces pricing-experiment candidates with confidence intervals; revenue lead approves discrete price changes. Never autonomous on production prices.
Cold call dialers with AI script + auto-record
Why it tempts you
Looks like cheap pipeline
Concrete failures
regulation TCPA + FCC ruling (see A3) -> AI voice without express consent is illegal. Two-party-consent state wiretap laws (CA, FL, IL, MA, MD, MT, NH, PA, WA) -> auto-recording without disclosure exposes private rights of action.
Right split
AI for opted-in re-engagement and inbound qualification only. All outbound voice still requires consent.
Investor update drafting without exec review
Why it tempts you
Monthly cadence, structured data
Concrete failures
judgment, low-frequency-high-stakes, regulation (securities) "Forward-looking statements" by LLM that don't accurately represent material facts create securities exposure. Investors notice tone shifts; auto-drafted updates damage trust.
Right split
AI compiles metrics deck and drafts narrative; CEO/CFO rewrite tone and approve.
Procurement decisions / RFP scoring
Why it tempts you
RFP scoring matrices feel objective
Concrete failures
judgment, vendor-lock, low-frequency-high-stakes LLM under-weights soft factors (vendor stability, cultural fit, post-sale support). Public-sector procurement frequently restricts AI in award decisions.
Right split
AI normalizes RFP responses, flags gaps and red flags; procurement committee scores and decides.
Lease term negotiation
Why it tempts you
Lease docs are templated
Concrete failures
judgment, relationship-dependent LL/T markets are highly local; what looks like a standard CAM provision is wildly different in Manhattan vs. suburban Atlanta. Concessions are negotiated based on relationships and market intel an LLM lacks.
Right split
AI compares draft to firm playbook + market comps; broker negotiates and partner signs.
Customer health scores driving auto-actions
Why it tempts you
Predict churn before it happens
Concrete failures
data-quality, judgment Stale CRM usage data + reverse causation generate false-positive churn flags; "save" outreach actually triggers reflection -> churn. High-value happy customers flagged red because they switched to a new champion (no login = "at-risk").
Right split
AI flags at-risk + suggests intervention; CSM owns the conversation and decides actions.
Auto-clawback / chargeback retry
Why it tempts you
Looks like recovered revenue
Concrete failures
regulation, cost Repeated retries on declined cards can trigger card-issuer flags + Stripe risk-score penalties.
Right split
AI determines optimal retry windows within Stripe Smart Retries policy; manual review on high-value or repeat failures.
Full code commits without human review (agent shipping to main)
Why it tempts you
Cursor / Devin / Copilot Agent feel close
Concrete failures
judgment, low-frequency-high-stakes, data-quality Cursor own incident — its support bot fabricated company policy (April 2025) (https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/). GitClear research 2024 found AI-assisted commits show 41% increase in code churn (rewrites within 2 weeks).
Right split
AI drafts PR; human reviews and merges. Auto-merge only for explicit allow-list (dependabot, doc typos, generated SDKs).
Auto-rollback of production deploys via AI
Why it tempts you
Faster MTTR
Concrete failures
judgment, low-frequency-high-stakes False-positive metric spikes (CDN cache flush, marketing campaign) trigger rollback during a legitimate launch -> revenue impact.
Right split
AI suggests rollback with evidence + page on-call; engineer executes after 1-look.
AI security incident response (autonomous blocking)
Why it tempts you
SOAR vendors pitch this
Concrete failures
judgment, change-management Autonomous IP blocks knock real customers off; auto-disabling user accounts on misclassified phishing triggers outage tickets.
Right split
AI enriches alerts + drafts containment plan; analyst executes. Auto-block only narrow, high-confidence rules (known-bad IPs).
AI listing description without fair-housing audit (real estate)
Why it tempts you
Saves agent hours
Concrete failures
regulation (Fair Housing Act) LLMs slip in "safe neighborhood," "family-friendly," "walking distance to church" — Fair Housing red flags. HUD complaints lead to brokerage fines + agent license risk.
Right split
AI drafts + rule-based scanner flags protected-class language; broker compliance reviews before MLS publish.
Auto-pricing of insurance / loan products
Why it tempts you
ML pricing is the industry's whole thing
Concrete failures
regulation (Fair Lending, anti-discrimination), data-quality CFPB and state DOIs require demonstrable non-discrimination; proxy variables (zip code) trigger disparate impact. Colorado SB21-169 specifically targets insurance algorithmic discrimination.
Right split
Actuaries + compliance own pricing; AI used for risk scoring within audited model framework.
Autonomous freight booking on load boards
Why it tempts you
Speed wins loads
Concrete failures
data-quality, judgment, fraud Double-brokering fraud rampant in 2024-25 — autonomous booking without carrier vetting transfers freight to bad actors. TIA tracking double-brokering losses exceeding $700M/yr industry-wide.
Right split
AI scores and surfaces top carriers + drafts rate confirmation; broker calls and books. Carrier identity verified through Highway/Carrier Assure (real-time MC# + identity check).
AI auto-prescribing or refilling Rx (without provider)
Why it tempts you
Cuts MA call burden
Concrete failures
regulation, low-frequency-high-stakes DEA + state pharmacy boards: only licensed prescribers can authorize. Repeated mis-refills (controlled substances) trigger DEA review + license risk.
Right split
AI queues refill requests with chart-pull summary; provider clicks approve/deny. Always.
Auto-payroll changes
Why it tempts you
Bonus calcs, comp adjustments
Concrete failures
regulation, judgment, low-frequency-high-stakes Wage-and-hour errors compound and back-pay claims accrue; California PAGA penalties are $200+/employee/pay period. One mis-coded raise propagates to bonus, equity refresh, severance baseline.
Right split
AI suggests + drafts; HR/finance approve before payroll cycle close. Audit log mandatory.
AI handling tenant lease-renewal negotiations
Why it tempts you
Renewal cadence is predictable
Concrete failures
relationship-dependent, regulation (rent control) Rent-stabilized markets (NYC, LA, SF, Oregon, statewide rent caps) bar certain increases — LLMs don't know which units in a portfolio are covered. Tenant relations damage when renewal feels transactional.
Right split
AI calculates market + suggests offer; PM personally communicates with tenants.
AI tenant screening (full auto-decline)
Why it tempts you
Volume of applications
Concrete failures
regulation (FCRA, FHA, state ban-the-box) FHA disparate impact on race + source-of-income protected classes in many states. FCRA requires adverse action notice with specifics — LLM can't cite its reason.
Right split
AI extracts and verifies (income docs, prior landlord refs); leasing agent makes go/no-go and sends compliant adverse action notices.
AI auto-quoting on rate-regulated services (healthcare prices, utilities)
Why it tempts you
Patient asks "what will this cost
Concrete failures
regulation (No Surprises Act, hospital price transparency rule), data-quality Federal hospital price transparency rule + state Good Faith Estimate rules require accuracy; misquoting can trigger HHS penalties and patient billing disputes.
Right split
AI pulls payer-specific rates + patient benefits; billing staff issue formal GFE.
AI receptionist for medical practice without HIPAA controls
Why it tempts you
Front desk is expensive
Concrete failures
regulation (HIPAA), privacy Many off-the-shelf voice AI vendors don't sign BAAs; transcripts of PHI passing through ungoverned LLM = breach. OCR breach notification + civil penalties.
Right split
AI receptionist with signed BAA, audit-logged, minimum-necessary PHI. Scheduling and FAQ only — never disclose chart info via untrained workflow.
Auto-drafted clinical notes posted directly to EHR
Why it tempts you
Reduces note-taking burnout
Concrete failures
regulation, judgment, low-frequency-high-stakes Hallucinated histories, transcription errors of medication doses (mg vs mcg) — direct patient safety risk. JAMA studies on AI scribe accuracy show clinically meaningful errors in 7-12% of notes.
Right split
AI scribe drafts; clinician reviews and signs every note (legal record).
Lead scoring with no human override
Why it tempts you
Reps want clarity
Concrete failures
data-quality, judgment Model drift; bad CRM hygiene poisons training data; sample-bias creates self-fulfilling prophecies.
Right split
AI suggests priority; rep accepts/rejects with reason; reasons feed retrain.
Generative AI inside customer-facing chatbots without prompt-injection defenses
Why it tempts you
ChatGPT-on-your-site
Concrete failures
data-quality, judgment, brand Chevrolet of Watsonville chatbot tricked into "agreeing" to sell a $76K Tahoe for $1 in November 2023 via the "Bakke Method" prompt injection (https://www.upworthy.com/chevy-chatbot-gone-wrong-ex1/, https://incidentdatabase.ai/cite/622/). OWASP top-1 risk for LLMs.
Right split
AI for FAQ from a controlled knowledge base; output guardrails (no committing to prices/policies); escalate to human for anything that touches money.
AI for SOC 2 / ISO control evidence gathering (auto-attest)
Why it tempts you
Audit prep is painful
Concrete failures
regulation, auditability Auditors require human attestation; auto-collected evidence that's mis-scoped won't satisfy CPA review and can constitute material misstatement.
Right split
AI gathers and indexes evidence; control owner attests. Tools like Vanta/Drata are this pattern done right.
AI auto-resolving employee benefits / leave requests
Why it tempts you
FMLA / STD intake is repetitive
Concrete failures
regulation (FMLA, ADA, state leave laws), privacy ADA interactive process requires individualized assessment; LLM rules engine cannot make that call.
Right split
AI handles forms intake + status comms; HR/leave specialist decides eligibility.
Real-estate AI showing-time scheduling that books seller property without listing-agent confirmation
Why it tempts you
Faster bookings
Concrete failures
change-management, relationship-dependent Tenant-occupied listings need scheduled access; auto-booking causes real-world conflict and broker complaints.
Right split
AI proposes windows; listing agent confirms; ShowingTime+ protocol stays in force.
AI auto-resolution of tenant disputes / habitability complaints
Why it tempts you
24/7 response
Concrete failures
regulation, low-frequency-high-stakes Habitability complaints have legal consequences (rent withholding, repair-and-deduct, retaliation claims).
Right split
AI triages and dispatches non-habitability work orders; PM owns any habitability or legal-implication complaint.
AI auto-quoting on construction change orders
Why it tempts you
CO disputes consume PM time
Concrete failures
judgment, contract-specific, low-frequency-high-stakes Each contract has different markup, time-impact, and lien-waiver rules; an LLM-issued CO can waive rights inadvertently.
Right split
AI drafts CO from field notes + photos; PM reviews and signs.
AI auto-scoring of customer service agent quality
Why it tempts you
100% of calls scored
Concrete failures
judgment, regulation, change-management Disparate impact on agents with accents or ESL backgrounds; union grievance risk; California PAGA exposure.
Right split
AI surfaces flag-worthy calls; QA team reviews; coaching conversations are human.
AI auto-generation of educational / training materials posted as company SOPs
Why it tempts you
SOPs are stale; LLMs fill the gap
Concrete failures
data-quality, judgment LLM SOPs include hallucinated tool names, deprecated steps; employees follow them and break things.
Right split
AI drafts SOPs from video + observed workflows; SME reviews and signs. Versioned in a wiki, with last-reviewed-by stamp.
Auto-drafted regulatory filings (BOI / 5500 / 5471 / VAT)
Why it tempts you
Forms are templated
Concrete failures
regulation, low-frequency-high-stakes FinCEN BOI penalty $591/day for inaccurate filings (https://www.fincen.gov/boi). IRS form 5471 penalty $10K per missed form; LLMs often misclassify foreign-entity types (CFC vs PFIC vs check-the-box election).
Right split
AI prepares draft + cites the rule it relied on; licensed pro reviews and signs.
Sales-call coaching with auto-PIP triggers
Why it tempts you
Conversation intelligence vendors pitch this
Concrete failures
judgment, regulation, change-management Auto-PIP based on call scores creates wrongful-termination exposure if disparate impact on protected classes. California, Illinois, Colorado AI Acts may classify this as workforce automated decision tool.
Right split
AI surfaces coaching themes per rep + per cohort; manager decides on PIP.
AI agent that takes actions in financial accounts (banks, brokerages)
Why it tempts you
Pay this bill, transfer $X
Concrete failures
regulation (KYC, Reg E, Reg D), low-frequency-high-stakes Plaid + Mercury TOS often restrict programmatic account actions; bad-actor scenarios trigger account freeze. Reg E liability for unauthorized transfers when AI scope is misconfigured.
Right split
AI drafts payment batches; CFO/owner approves in bank UI. Never store production credentials in LLM-accessible scope.
AI editing financial close adjusting journal entries
Why it tempts you
Month-end speed
Concrete failures
regulation, data-quality, auditability SOX controls for public co.; AICPA AU-C 240 fraud risk for private co. — auditors will reject AI-only JE workflow.
Right split
AI drafts + memo-explains; controller approves; CPA reviews.
AI auto-responding to subpoenas / legal holds
Why it tempts you
Templated process
Concrete failures
regulation, low-frequency-high-stakes Mishandled legal hold = spoliation sanctions, adverse inference instructions.
Right split
AI flags inbound + maps custodians; counsel manages response and preservation.
AI auto-applying for grants / RFPs
Why it tempts you
Volume play
Concrete failures
judgment, relationship-dependent, low-frequency-high-stakes Wrong eligibility claims = debarment risk; some grants require certification by signing officer. "AI-generated proposals" detection at some agencies leads to disqualification.
Right split
AI drafts proposal from past wins + grant guidelines; principal reviews and signs.
Driver dispatch optimization (full autonomy)
Why it tempts you
Big efficiency claim from vendors
Concrete failures
change-management, judgment, low-frequency-high-stakes Auto-dispatch ignores driver knowledge of difficult addresses, restaurant timing, weather; humans game it ("strategic deafness" to long-haul assignments). DOT HOS compliance: optimizer that pushes a driver into HOS violation = federal violation.
Right split
AI optimizer surfaces an ordered plan respecting HOS + driver constraints; dispatcher reviews and pushes.
AI agent committing trades, options, or DeFi transactions
Why it tempts you
24/7 markets
Concrete failures
regulation (SEC, CFTC), low-frequency-high-stakes FINRA Reg BI on suitability; market-manipulation exposure if LLM tries pump-and-dump tactics.
Right split
AI surfaces ideas + drafts orders; trader executes within risk limits.
Auto-publishing AI-generated articles to brand blog without review
Why it tempts you
SEO at scale
Concrete failures
regulation (FTC, plagiarism), data-quality, brand CNET's 77 AI-generated finance articles required 41 corrections in 2023; reputation damage. Google's March 2024 spam policy update targets scaled content abuse; deindexing risk.
Right split
AI drafts + structured outline; editor reviews, fact-checks, and signs.
AI builds "AI" features that don't actually use AI
Why it tempts you
AI-powered" sells
Concrete failures
regulation (FTC AI-washing) Builder.ai collapse May 2025 — used to be Engineer.ai accused of using humans behind "AI"; bankruptcy with $30M+ owed Microsoft (https://www.theregister.com/2025/05/21/builderai_insolvency/). FTC Operation AI Comply explicitly targets "machine learning when they relied on manual processes" (https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes).
Right split
If you claim AI, your stack must actually use AI; document substantiation; legal review of all AI claims.
Auto-deleting / archiving customer data on a schedule (no human review)
Why it tempts you
Data minimization compliance
Concrete failures
regulation (litigation hold, IRS / GAAP retention) Active legal hold + auto-delete = spoliation sanctions. Tax/financial retention requires 7-year hold; auto-deleting invoices at 3yr = IRS issue.
Right split
AI drafts retention policy and proposes records for deletion; legal/compliance approves batches.
AI customer ticket auto-close based on inactivity sentiment
Why it tempts you
Cleaner queue
Concrete failures
judgment, data-quality, change-management "Auto-closed for no response" creates measurable churn lift when applied to genuine issues; CS metrics game themselves.
Right split
AI nudges customer for response; agent decides to close after 2 nudges with no reply. Reopen on any customer reply.
Agent capability matrix — May 2026
decay: 3moWhat actually ships in production today. Reliability scores carry a 1.3× haircut versus vendor benchmarks. A demo is not a deployment.
Reality check before anything else
MIT's State of AI in Business 2025 found 95% of corporate GenAI pilots delivered zero measurable P&L impact across 300 deployments, 153 leader surveys, and 52 exec interviews. Fortune / MIT Gartner separately predicts >40% of agentic AI projects will be canceled by end of 2027. Gartner Same MIT data: purchased/partnered AI succeeds ~2× the rate of internal builds. Default audit recommendation: buy unless we have a defensible reason to build.
| Capability | SOTA (May 2026) | Rel | HITL | Cost/run | Top failure modes |
|---|---|---|---|---|---|
| Browser automation | Stagehand v3 + Browserbase + Claude Sonnet 4.5/4.6. WebVoyager near-saturated 88–89% bench; WebArena 68–74% leaderboard | 3/5 | Yes (writes) | $0.05–$0.50 | Site A/B drift; auth/captcha; injection via page text Unit42 |
| Code generation | Claude Opus 4.5/4.6/4.7 (80.9–87.6% SWE-bench Verified) & GPT-5/5.5 (74.9–88.7%) swe-bench. SWE-bench Pro top score 46% morphllm | 3–4/5 bounded, 2/5 long | Yes (PR) | $0.50–$5/task | Refactors >100k LOC; hallucinated APIs; Devin Railway hallucination Cognition |
| Email triage/draft | Superhuman, Shortwave, Lindy, native Gmail (Gemini) | 3/5 draft | Yes (send) | $9–$30/seat/mo | Tone drift; cross-thread hallucination; sensitive replies |
| Calendar scheduling | Reclaim, Motion, Cal.com AI. Clockwise sunsetting Mar 2026 Reclaim | 2/5 unsupervised | Yes (external) | $8–$19/seat/mo | Phantom events; over-defragmenting; cross-tz |
| Document / sheet | Claude + Skills (xlsx/docx/pptx); GPT-5 + Code Interpreter; Gumloop for batch | 3–4/5 | Yes (finance) | $0.10–$2/doc | Drops formulas; numeric vs text columns; XLSX >50 MB |
| CRM ops | HubSpot Breeze, Salesforce Agentforce, Clay, Lindy. HubSpot: "Customer Agent most production-ready" SMM | 3/5 | Yes (outbound) | $0.05–$0.50/record | Garbage-in/out; stale enrichment; dup rules over-fire |
| Voice agents | Retell ~600ms latency; Vapi ~500ms variable; ElevenLabs Agents Retell | 3/5 | Yes (escalate) | $0.05–$0.30/min | Hallucinated numbers; barge-in; TCPA exposure on outbound |
| Support deflection | Intercom Fin ~51% avg resolution Swifteq; Zendesk AI ~38% deflection. Klarna walked back AI-first PromptLayer | 3–4/5 mature KB, 2/5 day-one | Always tier 2+ | $0.50–$2/resolution | KB drift; "deflection theater"; brand voice |
| RAG / research | Claude Sonnet 4.5 + pgvector; GPT-5 file search; Perplexity API. Naive RAG retrieval fails ~40% of time lushbinary | 3/5 | Yes for decisions | $0.01–$0.50/q | Retrieval noise; conflated entities; confident-wrong |
| Data extraction | Claude 4.5 97–98% field acc; GPT-4V 95–96.5%; Gemini 3 93.8–95.8% tokenmix | 4/5 typed, 3/5 scanned | Yes >$threshold | $0.005–$0.10/doc | Layout edges; multi-page; locale decimals |
| Workflow orchestration | LangGraph (durable state); Temporal ($5B, OpenAI integration Sep 2025 InfoQ). OpenAI Agents SDK lacks checkpointing LangChain | 3–4/5 | Yes (branches) | $0.05–$5/run | State loss on restart; idempotency; 10+ step debug |
| Long-horizon | Claude Sonnet 4.5 "30-hr coding" claim; Gaia2 pass@1 only 42% Gaia2 | 2/5 | Checkpoints | $10–$100/day | Compounding error; 95%/step × 50 = 7% end-to-end medium |
| Computer-use (Claude/CUA) | Claude Mythos 79.6% OSWorld-Verified; OSWorld-Human shows efficiency lag arxiv | 2–3/5 | Yes (writes) | $1–$5/hr | OS dialogs; multi-monitor; UI-text injection |
Vendor & platform map — honest
decay: 3moWhat each platform is actually good at, where it breaks, and when to recommend / when to avoid. Skip the marketing pages.
Anthropic API + Claude Agent SDK
- Good at
- Claude Opus 4.5/4.6/4.7 leads SWE-bench Verified 80.9–87.6%; Sonnet 4.5 leads τ-bench Airline at 0.700. Agent SDK exposes tools, hooks, MCP, subagents. Skills for xlsx/docx/pptx/pdf.
- Breaks in prod
- Computer Use beta-ish, token-heavy. Opus pricey for high-volume routing.
- Lock-in
- Medium. MCP portable; Skills less so.
- Pricing landmines
- Extended-thinking multiplies cost. Cache aggressively — published patterns assume cache hits.
- Recommend when
- Coding, tool-heavy multi-step, doc extraction, long-context (200k–1M). Default for serious agent builds.
- Avoid when
- Native multi-vendor TTS/voice; sub-Haiku-only budgets.
OpenAI API + Agents SDK + AgentKit
- Good at
- GPT-5/5.5 raw capability; built-in `responses` with tools & file search; ChatGPT Agent (Operator absorbed July 2025). Temporal integration Sep 2025 for durability.
- Breaks in prod
- SDK lacks LangGraph-style checkpointing — HITL-pause flows need custom infra. Tool-call latency reports.
- Lock-in
- Medium–high. Responses API + AgentKit are OpenAI-specific.
- Pricing landmines
- Reasoning tokens; file search per-GB; Code Interpreter per-session. Caching exists but easy to miss.
- Recommend when
- Multimodal heavy; one-vendor preference; fine with OpenAI primitives.
- Avoid when
- On-prem; durable long-running state; vendor independence.
LangGraph + LangSmith
- Good at
- Production leader for stateful, branching, HITL-pauseable agent graphs. LangSmith traces are the most mature OSS option.
- Breaks in prod
- LangChain core is bloated — most teams now use only LangGraph + raw provider SDK.
- Lock-in
- Low for LangGraph; high for LangSmith cloud.
- Pricing landmines
- LangSmith trace volume; cloud deploy minimums.
- Recommend when
- Python team needs durable, stateful agents with HITL gates.
- Avoid when
- A managed product (Lindy, Gumloop, Zapier) covers it.
CrewAI
- Good at
- Quick demos with role-playing agents; rising mindshare.
- Breaks in prod
- Manager-worker process executes sequentially despite docs; agents skip tasks, fabricate IDs, return made-up results. Practitioner reports: '5–10 min/run, hours debugging silent failures.'
- Lock-in
- Low (Python).
- Pricing landmines
- High token use from multi-agent chatter.
- Recommend when
- Demos and hackathons.
- Avoid when
- Production unless you wrap heavily.
n8n (self-host + cloud)
- Good at
- v2.0 (Dec 2025) added isolated code execution, RBAC, 70+ AI nodes with LangChain integration. 80–90% cheaper than Zapier on high-volume flows. Only one with prod-grade self-host.
- Breaks in prod
- Steeper learning curve. Self-host means you operate HA/backups. Some integrations less polished than Zapier.
- Lock-in
- Low if self-hosted. AGPL is a gotcha for closed-source SaaS resale.
- Pricing landmines
- Cloud tier execution caps. AGPL gotcha for productized services.
- Recommend when
- Dev-led team, data-residency, agency builds for clients.
- Avoid when
- Non-tech solo founder; point-and-click only.
Zapier (incl. Zapier Agents)
- Good at
- Largest connector library (~8,500), polished templates, non-tech accessible. SOC 2 in flight on Agents (Dec 2025). Agents revamped May 2025 — focus moved from chat to automation.
- Breaks in prod
- Trustpilot 1.4 with surprise-billing complaints. AI builder fine for prototyping, not bulletproof prod. Linear workflows can't elegantly do branching/parallel. Zapier Central sunset.
- Lock-in
- High — workflows non-exportable.
- Pricing landmines
- Task billing; retries count. Multi-step AI Zaps compound.
- Recommend when
- Non-tech founder, low-volume, breadth of integrations dominant.
- Avoid when
- >10k tasks/mo; branching/parallel logic; self-host.
Make.com
- Good at
- Visual canvas with branching/parallel; SMB pricing; AI Agents added Oct 2025. Sits between Zapier (linear) and n8n (code).
- Breaks in prod
- Operations billing model is complex; harder to estimate. Some AI nodes thin.
- Lock-in
- Medium–high.
- Pricing landmines
- Operations metering on iterators (1000-row loop = 1000 ops).
- Recommend when
- SMB needing branching with a visual UI.
- Avoid when
- Self-host required.
Lindy.ai
- Good at
- Template-first AI-native agent creation; G2 4.9 (168+ reviews). Handles ambiguous tasks better than deterministic-workflow tools.
- Breaks in prod
- Closed system; debugging limited; integration breadth narrower than Zapier/n8n.
- Lock-in
- High.
- Pricing landmines
- Per-task + per-agent tiers stack.
- Recommend when
- Solo operator / agency owner needs assistant-style agents fast.
- Avoid when
- Complex DAGs or custom code paths.
Gumloop
- Good at
- Visual data pipelines purpose-built for PDFs, sheets, scrape, doc transform at scale. Top vendor score (84) in Zapier-alt comparisons.
- Breaks in prod
- Deterministic flow model less suited to ambiguous chat/email tasks.
- Lock-in
- Medium–high.
- Pricing landmines
- Per-run credits; large batches expensive.
- Recommend when
- Data extraction / batch document workflows.
- Avoid when
- Conversational agents primary.
Relevance AI
- Good at
- Agent-as-employee abstraction; sales/RevOps tooling; managed.
- Breaks in prod
- Less ecosystem than Lindy/Gumloop.
- Lock-in
- High (closed runtime).
- Pricing landmines
- Per-credit, opaque at high volume.
- Recommend when
- Sales/SDR team wants pre-built personas fast.
- Avoid when
- Need transparency or migration optionality.
Retool (+ Retool Agents)
- Good at
- Internal-tool UI + workflows + AI in one place. Strong RBAC. Best choice when agents need a human-facing dashboard.
- Breaks in prod
- Agent layer less mature than dedicated platforms; Retool-style debug pain.
- Lock-in
- Very high (custom DSL/state).
- Pricing landmines
- Per-seat scales painfully; Retool DB / workflows are stacked add-ons.
- Recommend when
- Internal ops dashboard with agent assist.
- Avoid when
- Public/customer-facing; cost-sensitive.
Airtable + Cobuilder
- Good at
- Cobuilder spins up bases from spec. AI fields for enrichment. Ubiquitous SMB adoption.
- Breaks in prod
- Scaling walls at ~100k records / complex automations. AI features paywalled.
- Lock-in
- High data model.
- Pricing landmines
- Per-user seats; AI credit add-ons.
- Recommend when
- SMB CRM-lite, content ops, light agents on top.
- Avoid when
- True backend / app-level data.
Supabase
- Good at
- Postgres + Auth + Storage + Edge + pgvector. Real-time. Default agent backend.
- Breaks in prod
- Free tier pauses; RLS is power-user only.
- Lock-in
- Low (Postgres).
- Pricing landmines
- Egress; compute add-ons.
- Recommend when
- Any production agent backend that needs DB+auth.
- Avoid when
- Fully managed enterprise (Snowflake-class).
Browserbase + Stagehand
- Good at
- Stagehand v3 is 44%+ faster than v2 via direct CDP. Self-healing selectors; iframe/shadow DOM. Browserbase cloud handles captchas, session replay, agent identity.
- Breaks in prod
- AI-resolved selectors add latency and LLM cost. Non-deterministic debug.
- Lock-in
- Low for Stagehand OSS; high for Browserbase cloud.
- Pricing landmines
- Per-session minutes + LLM tokens.
- Recommend when
- Any agent doing real browser tasks across changing sites.
- Avoid when
- Single stable site you own (use Playwright).
Playwright / Puppeteer
- Good at
- Cheapest, fastest, deterministic. Existing test-automation expertise transfers.
- Breaks in prod
- Brittle — redesigns break selectors instantly. No 'AI healing.'
- Lock-in
- None.
- Pricing landmines
- Compute only.
- Recommend when
- Stable site + scheduled job.
- Avoid when
- Sites you don't own or change often.
Firecrawl
- Good at
- LLM-ready markdown; thousands of pages/min mid-tier. AI-driven navigation. Single API for scrape/crawl.
- Breaks in prod
- No native scheduling — bring your own. Less control on dynamic anti-bot.
- Lock-in
- Low.
- Pricing landmines
- Per-page credits; JS-render multiplier.
- Recommend when
- RAG ingestion, research-agent feeds.
- Avoid when
- Need scheduling / queue / Actor marketplace.
Apify
- Good at
- Scheduler + queue + retries built-in. Actor marketplace. Burst workloads (thousands parallel).
- Breaks in prod
- Heavier learning curve; container overhead.
- Lock-in
- Medium.
- Pricing landmines
- Compute units + proxy traffic.
- Recommend when
- Production scrape pipelines, e-com catalogs.
- Avoid when
- Just need 'LLM-ready text' — use Firecrawl.
Vector DBs (Pinecone / Qdrant / Weaviate / pgvector)
- Good at
- At 10M vectors: pgvector ~$45/mo, Qdrant ~$65, Pinecone serverless ~$70, Weaviate ~$135. At 100M: pgvector/Milvus <$100 vs Pinecone $700+. Qdrant 1840 QPS on 1M vectors.
- Breaks in prod
- Pinecone serverless cold-start latency; Weaviate ops if self-hosted.
- Lock-in
- Low for pgvector/Qdrant; high for Pinecone managed features.
- Pricing landmines
- See cost matrix — vector count + QPS is the metric, not storage.
- Recommend when
- <5M + Postgres shop → pgvector. <100ms P99 no-ops → Pinecone. Hybrid keyword+vector → Weaviate. High-QPS OSS → Qdrant.
- Avoid when
- Wrong tier for scale — switching mid-deploy is painful.
Voice — Retell / Vapi / ElevenLabs Agents
- Good at
- Retell ~600ms latency, SOC2/HIPAA/GDPR out of box. ElevenLabs voice quality leader; IBM watsonx partnership Mar 2026. Vapi flexible if you tune.
- Breaks in prod
- Vapi sub-500ms is variable; carrier latency adds 200ms+ to all. TTS misreads on numbers.
- Lock-in
- Medium per platform.
- Pricing landmines
- Per-minute talk time + per-LLM-call. PSTN trunking extra.
- Recommend when
- Regulated → Retell. Brand voice → ElevenLabs. Custom tuning + voice eng → Vapi.
- Avoid when
- Outbound cold sales (TCPA exposure).
Integration APIs — Paragon / Merge / Nango
- Good at
- Nango: free 5k calls/mo, $249 for 50k single-category — most flexible OAuth + sync logic you own. Merge: unified API (HRIS/ATS/Accounting/CRM/ticketing). Paragon: enterprise embedded.
- Breaks in prod
- Paragon pricing opaque (no clear metric on page). Merge: less control on edge sync logic.
- Lock-in
- Low Nango, medium Merge, high Paragon.
- Pricing landmines
- Per-record/calls for Nango/Merge; sales-led for Paragon.
- Recommend when
- Startup → Nango. B2B SaaS embedded (many integrations) → Merge. Enterprise sales motion → Paragon.
- Avoid when
- Wrong tool for stage.
Recommended stacks — by use case
decay: 3moPractical defaults. Pick by the closest-fit use case, not the most exciting vendor.
- No-code MVP, non-technical founder. Lindy (or Zapier Agents) + Gmail/Calendar/CRM connectors + Claude Sonnet 4.5 backend. Fastest time-to-first-agent. Accept lock-in for speed — MIT: bought beats built ~2:1.
- SMB automation, agency-owner-no-devs. Make.com OR n8n cloud + Claude/OpenAI + Airtable/Supabase + Firecrawl/Apify. Visual canvas + branching + AI nodes. n8n if self-host or data-residency matters; Make if Make is what they know.
- Internal ops dashboard. Retool (UI + workflows + agents) + Supabase + Claude API. Retool is the unique combo of UI builder + agent runtime. Worth the lock-in for ops portals.
- Browser-heavy. Stagehand v3 on Browserbase + Claude Sonnet 4.5 + Temporal (durability) + Supabase. Self-healing selectors are the only thing that survives DOM drift at scale. Temporal carries long browser sessions across failures.
- Document-heavy. Gumloop (batch) OR Claude API + Skills (xlsx/pdf/docx) orchestrated by LangGraph, persisted in Postgres. Claude leads invoice extraction (97–98%) and produces valid JSON 100% of cases.
- Sales/support automation. Intercom Fin OR Zendesk AI Agents (deflection) + Retell voice + Clay/HubSpot for outbound enrichment + Merge unified API to sync CRMs. Buy the deflection layer (Fin > 50% avg resolution beats most internal builds). Humans on tier 2. Learn Klarna's lesson — don't go AI-first 100%.
- Productized service (build once, sell many). n8n self-hosted (watch AGPL) OR LangGraph + OpenAI/Anthropic + Supabase + Browserbase OR white-label Lindy/Make per-tenant. LangGraph stack is most defensible/portable; managed stack is fastest.
Failure modes & cross-cutting heuristics
decay: 6moThe 15 most common reasons agent deployments fail in 2025-2026, with citations. Plus the heuristics our audit agent applies.
Top 15 failure modes
- Eval gap. Naive RAG retrieval fails ~40% in prod lushbinary; 95% of corp pilots produce zero P&L MIT.
- Long-horizon compounding. 95%/step × 50 steps = 7% success medium. METR 50%-reliability horizon: tens of minutes, not days METR. Gaia2 ceiling 42% pass@1.
- pass^k collapse.τ-bench retail pass@1 < 50% but pass^8 < 25% Sierra— your demo isn't your 8th-run reality.
- Prompt injection. CVE-2025-59944 (Cursor → RCE); $250k bank-assistant fraud (Jun 2025); AI worm Feb 2025 mayhem. Multi-hop indirect attacks +70% YoY SQ.
- Brittle browser DOMs. Even Stagehand self-healing carries LLM-resolution cost per repair. Princeton "Illusion of Progress" shows agents fail on real-world long-tail sites arxiv.
- Hallucinated tool calls. CrewAI agents fabricate IDs / skip tasks Popelka; Devin hallucinated Railway features for a day+ Cognition.
- Context cost spiral. Long-running agents accumulate context; computer-use screenshots especially heavy. Budgets blow on retries.
- Permissioning hell. Real deployments hit OAuth scope sprawl across Gmail/Cal/CRM. Anthropic guidance: deny-all-by-default, allowlist per subagent Anthropic.
- Vendor pricing surprises. Zapier Trustpilot 1.4 Startupowl. Make ops billing on iterators. Pinecone $700+ at 100M vs pgvector <$100 LeanOps.
- HITL fatigue → rubber-stamping. Operators learn to approve without reading. Klarna CSAT drop → AI-first reversal CXD.
- Reward hacking / spec gaming. UC Berkeley CRDI: an automated scanning agent broke all 8 major agent benchmarks via reward hacking (Apr 2026) rapidclaw.
- Statelessness / restart loss. Agents commonly forget previous messages, lack retries, lose progress on reboot Particula. OpenAI Agents SDK lacks checkpointing.
- Fan-out that doesn't fan out. CrewAI sequential despite docs implying coordination TDS.
- Voice hallucinations & latency.<800ms required for natural feel; some platforms 3–4s. TTS misreads numbers (account IDs, dates); TCPA risk on outbound.
- Benchmark contamination. OpenAI dropped SWE-bench Verified — moved to SWE-bench Pro where top score is 46% vs 81% on Verified morphllm. Audit-agent rule: discount any single benchmark ~30% when projecting to a customer's domain.
Cross-cutting heuristics (the audit agent's rules)
System prompt — heuristics block
- Prefer buy over build by ~2:1. MIT data is clear.
- Quote a reliability haircut: divide vendor-claimed accuracy by ~1.3× for real-world distribution shift.
- Always quote pass^k or k-attempt reliability, not pass@1.
- Budget retries. 30% retry rate at $0.50/run = $0.15 hidden cost per task.
- Mandatory HITLfor: send (email/SMS), spend (>$X), legal/compliance, customer-visible writes.
- Pick durability infra first. LangGraph + Temporal (or Mistral Workflows pattern). Statelessness kills more deploys than model quality.
- Default vector DB: pgvector if you have Postgres.
- Default doc extraction: Claude 4.5/Opus. GPT-4o only on degraded scans.
- Voice in regulated industries: Retell or ElevenLabs. SOC2/HIPAA/GDPR out of box matters.
- Browser: Stagehand+Browserbase over Playwright unless target is owned and stable.
Per-role pain inventory
decay: 12moWho buys, what they hate, what they'll pay for. Each role maps to processes from §3.
Agency Owner (services, 5–50 people)
$120k–$400k+ owner drawHeadcount — SMB: 1 founder · Mid: 1–3 owners + operators · Enterprise: n/a — sells to large agencies
Top pains (their words)
- "Friday afternoon panic-deck reporting eats my week."[src]
- "Scope creep is silently destroying our margin per project."[SPI Research PSMB 2024]
- "Time entry compliance is a Friday email tax."[Workamajig agency surveys]
- "Proposals take 3 days when they should take 30 minutes."[PandaDoc benchmarks 2024]
- "Tracking utilization vs. signed SOWs is a stitched-spreadsheet nightmare."[SPI Research PSMB]
Weekly hours automatable
12–18 hrs founder time
Buying authority
Full sign-off on $0–$3k/mo on AmEx; founders are the buyer.
Existing band-aids
Templates in Notion, ClickUp, Workamajig. Most have tried 1-2 prior automation attempts that fizzled on adoption.
Director of Operations (SMB / Mid)
$120k–$180k base + bonusHeadcount — SMB: 1 · Mid: 1 + 1–3 IC ops · Enterprise: VP Ops + Director(s) + 5+ IC
Top pains (their words)
- "We use 9 tools that don't talk and I'm the human SQL query."[Asana Anatomy of Work 2024]
- "Weekly KPI rollups eat 4 hours every Monday."[Operator community forums]
- "Status updates between teams get dropped on handoff."[Atlassian State of Teams 2024]
- "Our spreadsheets are our database and nobody trusts them."[MetaPlane data-quality survey]
- "Vendor renewals auto-fire without negotiation."[Vendr State of SaaS 2024]
Weekly hours automatable
14–22 hrs
Buying authority
$0–$5k/mo direct; above goes to CFO/founder.
Existing band-aids
Zapier, Make, Notion automations, a few SOPs. Most are partially-functional.
CFO at 20–100 person company
$220k–$360k base + equityHeadcount — SMB: Fractional CFO + bookkeeper · Mid: 1 CFO + 2–5 finance ops · Enterprise: VP Finance + Controller + 10+ ICs
Top pains (their words)
- "Month-end close takes 12 days when it should take 5."[FloQast Close Benchmark 2024]
- "AR follow-ups consume 30%+ of AR specialists' time."[src]
- "Cash flow forecasts are last week's data by the time they're done."[AFP Cash Forecasting Survey]
- "Board pack assembly Friday before is a fire drill."[FENG community AMAs]
- "I can't tell which customers are profitable."[AICPA SMB benchmarking]
Weekly hours automatable
10–25 hrs across finance team
Buying authority
$5k–$50k/mo discretionary; gatekeeper for everyone else's AI spend.
Existing band-aids
QBO/NetSuite + FloQast/Numeric for close; Ramp/Brex for cards; the FP&A model is in Sheets.
RevOps Lead
$140k–$220k loadedHeadcount — SMB: 0–1 part-time · Mid: 1 + 1–2 ICs · Enterprise: Director + 5+ analysts
Top pains (their words)
- "Forecasting is a vibe check; the CRO doesn't trust the number."[Clari forecasting reports]
- "CRM hygiene drift kills every dashboard; nobody owns it."[Gong Pipeline Report 2024]
- "Lead routing rules are spaghetti and leads sit too long."[InsideSales / Velocify 5-minute rule research]
- "Pipeline reviews are slide-rebuilds, not strategy."[Gong State of Revenue 2024]
- "Sales notes never make it to CRM; reps hate it."[src]
Weekly hours automatable
15–25 hrs
Buying authority
$0–$3k/mo direct; CRO/CFO above.
Existing band-aids
HubSpot/Salesforce + Gong + Outreach + custom Zaps. CRM is the bottleneck.
Head of Customer Success
$160k–$250k loadedHeadcount — SMB: Founder + 1 CSM · Mid: Director + 3–8 CSMs · Enterprise: VP CS + Director(s) + 20+ CSMs
Top pains (their words)
- "Renewals sneak up; we start motion 30 days out when we need 90."[Gainsight 2024 State of CS]
- "QBRs eat a CSM's entire week each quarter."[CSM community forums]
- "Health scores are gut feel; churn surprises us."[Catalyst Customer Health benchmark]
- "Expansion always gets deprioritized for save-the-account fires."[Gainsight 2024]
- "Onboarding is ad-hoc; no two customers see the same playbook."[Rocketlane onboarding benchmarks]
Weekly hours automatable
18–30 hrs across the CS org
Buying authority
$1k–$10k/mo direct; CRO above for enterprise tools.
Existing band-aids
Gainsight/Catalyst/Vitally + Slack + spreadsheets. Most run rules-based health, not ML.
E-commerce Operator ($1–10M Shopify brand)
$80k–$300k owner drawHeadcount — SMB: Founder + 1–3 · Mid: Founder + 5–15 ops/CS/marketing · Enterprise: n/a
Top pains (their words)
- "Reviews pile up on Trustpilot and Amazon; responses are inconsistent."[Trustpilot benchmark 2024]
- "Returns processing is a manual customer-service grind."[Loop Returns benchmark]
- "Ad creative testing is slow; we never have enough variants."[AdCreative.ai benchmarks]
- "Customer service for order issues is 60% of tickets."[Gorgias / Zendesk e-comm benchmark]
- "Post-purchase email flows are out of date."[Klaviyo State of Email 2024]
Weekly hours automatable
20–35 hrs across team
Buying authority
Founder full sign-off on $0–$5k/mo on AmEx.
Existing band-aids
Shopify app stack (12+ tools), Klaviyo, Gorgias, Loop. Most over-tooled and under-integrated.
Solo SaaS Founder / Indie Hacker
$0–$300k (lumpy)Headcount — SMB: 1 · Mid: n/a · Enterprise: n/a
Top pains (their words)
- "Inbox drowns; important sales/press emails get missed."[Indie Hackers forums]
- "Support tickets at 50/day require my attention until I hire."[Indie Hackers / r/SaaS]
- "I'm the only one who knows how anything works."[MicroConf community]
- "Manual content distribution and repurposing is 8 hrs/week."[MicroConf surveys]
- "Time-to-bill is too long because I forget to invoice."[r/SaaS founder threads]
Weekly hours automatable
10–20 hrs founder time
Buying authority
Full sign-off on $0–$2k/mo.
Existing band-aids
Stripe + Notion + ChatGPT + ad-hoc Zaps. Building automations in spare time.
Practice Manager (medical / dental / legal / accounting)
$75k–$140kHeadcount — SMB: 1 + 2–5 admin · Mid: 1–2 + 6–20 admin · Enterprise: n/a (MSO instead)
Top pains (their words)
- "Patient/client intake is a paper-form graveyard."[AMA practice ops survey 2024]
- "Insurance verification and prior auth chases dominate the day."[src]
- "Claim denials are a recurring 8-15% revenue tax."[src]
- "Recall campaigns for return visits are inconsistent."[Practice Management Institute]
- "Referral coordination between practices is fax + phone."[AAPP practice surveys]
Weekly hours automatable
15–25 hrs across admin team
Buying authority
$0–$2k/mo direct; physician/partner above.
Existing band-aids
Tebra / Athena / Clio + paper. HIPAA / state bar / accounting rules constrain options.
Local Services Business Owner (HVAC, plumbing, landscaping)
$80k–$300k owner drawHeadcount — SMB: 1 + 2–6 techs · Mid: 1 + 8–20 techs · Enterprise: n/a — sells to franchises
Top pains (their words)
- "Leads sit for 30+ minutes before someone responds. Conversion craters."[src]
- "Techs forget to update job status; dispatch is guessing."[ServiceTitan customer benchmarks]
- "Estimates eat our weekends."[Buildxact contractor surveys]
- "Review requests are inconsistent post-job."[Birdeye/Podium benchmarks 2024]
- "Recurring service reminders rely on memory, not system."[ServiceTitan 2024]
Weekly hours automatable
10–20 hrs owner + dispatcher
Buying authority
Owner full sign-off on $0–$2k/mo.
Existing band-aids
ServiceTitan / Jobber / Housecall Pro. Most leave money on the table by not using the AI features.
Engineering Manager (50–500 person co)
$220k–$360k loadedHeadcount — SMB: 1 + 4–8 eng · Mid: 1 + 8–15 eng · Enterprise: Director + multiple EMs
Top pains (their words)
- "PR review is a senior-engineer bottleneck."[DORA 2024 State of DevOps]
- "Onboarding a new engineer takes 4–8 weeks to first material PR."[Stack Overflow Developer Survey]
- "On-call burns out senior engineers; runbooks are stale."[PagerDuty 2024 oncall report]
- "Flaky tests kill CI trust."[Buildkite reliability surveys]
- "Internal docs are stale; same questions get asked weekly."[Atlassian State of Teams 2024]
Weekly hours automatable
20–35 hrs across eng org
Buying authority
$1k–$5k/mo direct; VP Eng / CTO above.
Existing band-aids
GitHub Copilot + Cursor + manual PR review. Cody/Greptile in some shops.
Recruiting Operations / Talent Ops
$95k–$160k loadedHeadcount — SMB: 0–1 · Mid: 1 + 1–3 recruiters · Enterprise: Director + 5+ ICs
Top pains (their words)
- "Sourcing personalization at scale is impossible manually."[Gem 2024 Talent Acquisition]
- "Interview scheduling across panel + candidate burns 30 min/loop."[GoodTime benchmarks]
- "Candidate ghosting happens because reply SLAs slip."[src]
- "Reference checks are last-minute manual scrambles."[Crosschq surveys 2024]
- "ATS data is dirty; reporting is broken."[Ashby benchmarking 2024]
Weekly hours automatable
12–22 hrs
Buying authority
$500–$2k/mo direct; Head of Talent above.
Existing band-aids
Greenhouse/Ashby/Lever + Gem/SourceWhale + GoodTime. The stack is mature; integration is the gap.
Chief of Staff / Executive Assistant
$140k–$280k loadedHeadcount — SMB: 0–1 EA · Mid: 1 CoS + 1 EA · Enterprise: CoS + EA staff
Top pains (their words)
- "Founder's inbox is a fire-hose; important items get missed."[Founder community forums]
- "Board pack assembly is a Friday-night ritual."[CoS community AMAs]
- "Founder shows up to meetings cold without context."[EA Network surveys]
- "OKR/metric rollups across functions take a full day."[Operator Collective community]
- "Scheduling negotiation across exec calendars is a multi-day game."[GoodTime exec scheduling]
Weekly hours automatable
20–30 hrs
Buying authority
$0–$3k/mo direct on AmEx; founder sponsorship for larger.
Existing band-aids
Superhuman/Shortwave + Notion + Google Workspace + manual scripts.
Operator ↔ vendor glossary
decay: stableHow operators actually describe pain (left) vs. the tech category that solves it (right). Sales / discovery gold.
| Operator phrase | Capability | Example solution |
|---|---|---|
| "Leads sit too long before we touch them." | Inbound lead routing + speed-to-lead automation with SLA timer | HubSpot/Salesforce workflows + Default/Distribute + Slack-alerting agent that pings AE if no first-touch in 5 minutes[src] |
| "I'm copying and pasting between LinkedIn, Apollo, and the CRM all day." | Browser-based prospecting orchestration / sequencer enrichment | Clay or Apollo + n8n/Make + CRM write-back agent |
| "Sales notes never make it to CRM." | Call → CRM enrichment (transcript → field extraction) | Gong/Fathom/Granola + LLM extraction agent + Salesforce/HubSpot API write[src] |
| "I don't trust our CRM data." | CRM hygiene + dedup + enrichment automation | Syncari/Openprise or custom dbt + Clearbit/Apollo enrichment + nightly LLM dedup agent |
| "Forecasting is a vibe check." | Pipeline scoring + commit-call analytics | Clari/BoostUp or Salesforce + Gong forecast + LLM stage-validation agent that flags stale deals[src] |
| "I never know what my pipeline really looks like." | Real-time pipeline reporting + deal hygiene agent | Salesforce reports + Gong + Slack digest agent that posts MEDDPICC gaps daily |
| "My SDRs send the same email 80 times a day." | Personalized outbound at scale | Clay + Apollo + LLM personalization on top of Outreach/Salesloft sequences |
| "Discovery calls are inconsistent across reps." | Call coaching + question-coverage scoring | Gong/Chorus + LLM scorecard agent that flags missed MEDDPICC fields |
| "Quotes take days to get out the door." | CPQ + approval workflow automation | Salesforce CPQ / DealHub + Slack approval agent |
| "Contract redlines bounce around for weeks." | Contract lifecycle management with AI redline assistant | Ironclad/Lexion + LLM redline-comparison agent + DocuSign |
| "Renewals sneak up on me." | Renewal forecasting + auto-prompted CSM workflow | Gainsight/Catalyst + renewal-90/60/30 agent in Slack with talking points |
| "Our QBRs eat my whole week." | Auto-generated QBR decks from usage + CRM data | Mixpanel/Amplitude + CRM + LLM deck-builder agent (Google Slides API) |
| "Customer health is a gut feel." | Health scoring from product usage + support + sentiment | Gainsight/Vitally + Zendesk + Gong sentiment + LLM aggregator |
| "Onboarding is a mess." | Workflow orchestration with deadline timers + Slack nudges | Rocketlane/Arrows + Slack agent + checklist-tracking LLM |
| "Churn shows up out of nowhere." | Leading-indicator churn agent (usage drop + sentiment + support volume) | Catalyst/Vitally + Zendesk + LLM signal-aggregation agent |
| "Expansion never gets prioritized over saves." | Expansion-signal detection (feature adoption, seat growth) | Product analytics + CRM + LLM agent that surfaces expansion plays weekly |
| "Tickets pile up overnight." | 24/7 deflection agent + L1 triage | Intercom Fin / Zendesk AI / Decagon / custom RAG over docs + KB[src] |
| "We answer the same question 50 times." | KB-grounded chatbot with deflection | Zendesk AI / Intercom Fin / custom RAG (Pinecone + Claude) + Slack escalation |
| "Agents waste 10 minutes searching the wiki for every ticket." | Agent-side AI copilot with retrieval | Zendesk AI / Forethought / Cresta + internal KB embeddings |
| "Macros are out of date but nobody updates them." | Auto-generated response suggestions from resolved tickets | LLM mining past tickets → suggested macros pipeline |
| "Tier 1 escalates everything because they're scared to answer." | Confidence-scored answer suggestions + decision tree agent | Forethought or custom LLM with confidence threshold + human handoff |
| "We can't keep up after a product launch." | Surge handling with topic clustering + auto-FAQ generation | Topic-clustering LLM over inbound + auto-update KB + macro suggestion |
| "Following up on AR is killing me." | Automated dunning / collections agent | Upflow/Chaser/Versapay + email-personalization LLM + Slack escalation[src] |
| "Vendor invoices show up everywhere." | AP intake + OCR + 3-way match automation | Ramp Bill Pay / Bill.com / Tipalti + LLM extraction + ERP push |
| "Month-end close takes 12 days." | Close orchestration + recon agent | FloQast / Numeric + LLM JE-explanation agent |
| "Reconciliations are 80% of my month." | Auto-recon with exception-only review | BlackLine / Numeric + LLM variance explainer |
| "Budget vs. actuals takes a week to assemble." | Live FP&A dashboard with variance commentary | Cube / Pigment / Mosaic + LLM variance-narrative agent |
| "Expense reports are a nightmare." | Card-feed + receipt OCR + auto-categorization | Ramp / Brex + LLM policy-check agent |
| "Audit prep is a full-time job for two weeks." | Continuous controls monitoring + evidence collection | AuditBoard / Drata + evidence-gathering agent |
| "I can't tell which customers are profitable." | Unit economics dashboard with cost allocation | dbt + Mode/Hex + LLM commentary agent |
| "We keep dropping the ball on handoffs." | Cross-system state tracking + checklist agent | Linear/Asana + Slack agent + LLM workflow-state monitor with reminders |
| "We use 9 tools that don't talk." | iPaaS + custom integration layer | Workato / Tray / n8n + event bus + LLM mapping helper |
| "Spreadsheets are our database." | Lightweight workflow DB + form intake | Airtable / Smartsheet / Retool + migration agent |
| "Manual data entry between systems." | ETL + system-of-record sync | Fivetran / Hightouch + reverse-ETL + LLM field-mapping helper |
| "Status updates eat my Mondays." | Auto-generated status reports from Linear/Jira/Slack | Linear + Slack + LLM weekly-summary agent |
| "Standups are pointless." | Async standup bot with auto-extracted blockers | Geekbot / Range + LLM blocker-detection agent |
| "Two people own this, neither does it." | RACI tracking + accountable-owner agent | Notion DB + Slack reminder agent that escalates after N days |
| "We rebuild the same deck every Monday." | Templated reporting deck auto-generated from data | Google Slides API / Plus.ai + LLM commentary layer |
| "Nobody reads our docs." | RAG-powered Q&A over internal wiki | Notion/Confluence + Glean / Dust / custom RAG with Slack bot |
| "I'm drowning in Slack." | Smart Slack digest + thread summarization | Slack AI / Glean / custom agent with channel-priority logic |
| "Inbox zero is a myth." | Email triage agent with draft replies | Superhuman AI / Shortwave / Missive + LLM draft agent |
| "We can't find the latest version of anything." | Single-source-of-truth doc system + version intelligence | Notion + Glean federated search agent |
| "Customer requests get lost in DMs." | DM → ticket capture agent | Slack listener bot → Linear/Zendesk auto-create |
| "We never have data when leadership asks." | Self-serve analytics + LLM-over-warehouse | Hex / Mode / Julius + dbt + LLM SQL-generation layer |
| "Hiring is a black hole." | ATS automation + candidate-status visibility | Ashby/Greenhouse + Slack agent + LLM screening assist |
| "Recruiters ghost us." | Pipeline SLA tracking + candidate-response agent | Ashby + LLM personalization for outreach + nudge agent |
| "Performance reviews are theater." | Continuous-feedback platform + AI-summarized reviews | Lattice / 15Five + LLM that aggregates 1:1s, PRs, and Slack into draft review |
| "Onboarding new hires is ad-hoc." | Onboarding workflow agent + checklist tracking | Rippling / Sapling + Slack onboarding bot |
| "PTO requests bounce around email." | Self-serve HRIS workflows | Rippling / Gusto / Deel + Slack approval bot |
| "Comp benchmarking is a guessing game." | Comp data + leveling automation | Pave / Figures + LLM offer-letter generator |
| "Compliance is always a fire drill." | Continuous compliance automation | Drata / Vanta / Secureframe + evidence agent |
| "Legal needs to approve everything." | Self-serve legal playbook + AI redline | Ironclad / Spellbook + LLM playbook-routing agent |
| "NDAs take a week." | NDA auto-execution with playbook | Ironclad / LinkSquares + LLM clause-check |
| "Vendor security reviews kill deals." | Auto-completed security questionnaires from policy KB | Loopio / Responsive + RAG over policies |
| "Marketing requests pile up in JIRA." | Creative ops intake + auto-routing | Asana / Wrike + LLM brief-completeness agent |
| "Brief never matches what's delivered." | Structured brief intake with required-field validation | Form + LLM brief-quality scorer + approval workflow |
| "We can't attribute pipeline to campaigns." | Multi-touch attribution + LLM commentary | HubSpot / Bizible / Dreamdata + revenue-impact agent |
| "Content calendar is in three places." | Single content ops hub | Notion/Airtable + Asana + LLM repurpose agent |
| "SEO content takes forever to brief." | AI brief generator from SERP + competitor data | Ahrefs/SEMrush API + LLM brief-writer |
| "Newsletter is always late." | Auto-curated newsletter from RSS + KB + LLM editor | Beehiiv/Substack + LLM curation agent |
| "PagerDuty is keeping me awake." | Incident triage + auto-runbook execution | PagerDuty + Rootly / FireHydrant + LLM runbook agent |
| "Code review takes days." | AI PR reviewer + auto-summarization | GitHub + Greptile / CodeRabbit / Graphite + LLM reviewer |
| "Onboarding a new engineer takes a month." | AI codebase tour + setup automation | Cody / Cursor / internal RAG agent over repos |
| "IT tickets for password resets are 40% of volume." | Self-service IT bot | Moveworks / Atomicwork / custom Slack agent |
| "Shadow IT is everywhere." | SaaS discovery + spend management | Zylo / Torii / Vendr + LLM categorization |
| "Procurement is a forest." | Intake-to-pay workflow with policy routing | Zip / Tropic + LLM contract-summary agent |
| "Vendor renewals just happen." | Contract repository + renewal-alert agent | Vendr / Tropic + LLM negotiation-prep agent |
| "We don't know what we're paying for." | SaaS spend visibility + usage-based rightsizing | Zylo / Torii + usage-pull agent |
| "I'm the bottleneck for everything." | EA agent + delegation tracking | Lindy / Cora / custom agent for inbox, scheduling, follow-up |
| "Time tracking is hated." | Passive time-capture from calendar + tools | Reclaim / Memtime + LLM categorization |
| "My calendar is on fire." | AI scheduler with preference learning | Reclaim / Motion / Clockwise |
| "I have no visibility into what my team is actually doing." | Cross-tool activity dashboard + LLM weekly synth | Linear + GitHub + Slack + LLM exec digest |
| "Customers ask for the same thing in DMs." | DM intake → ticket + auto-FAQ | Slack/Telegram listener + LLM categorizer + Notion FAQ writer |
| "Insurance claims sit for weeks." (insurance ops)" | Claims triage + document extraction | Hyperscience / Sensible + claims-LLM agent + Guidewire write |
| "Field techs forget to update job status." (field services)" | Mobile-first job update + voice-note → CRM | ServiceTitan / Jobber + voice-transcription LLM |
| "Patient intake forms are a disaster." (healthcare ops)" | Pre-visit intake + EHR sync (HIPAA-scoped) | Phreesia / Tebra + form-LLM with PHI controls |
| "Construction RFIs take forever to answer." (construction)" | RFI auto-draft from drawings + spec | Procore + RAG over project docs + LLM draft agent |
| "Estimates eat our weekends." (trades / contracting)" | AI estimator from photos + scope notes | Buildxact / custom + vision-LLM takeoff agent |
| "Restaurant scheduling is chaos." (hospitality)" | AI shift scheduler with demand forecast | 7shifts / Homebase + forecasting agent |
| "Legal intake is a Google form graveyard." (law firms)" | Client-intake agent + matter creation | Clio + LLM intake bot + conflict-check agent |
Discovery question bank
decay: 12moThe intake script the audit agent runs. Conditional branches by function. Wrap in <sysprompt> in deployment.
System prompt block — intake agent
You are the AutomationAudit intake agent. For each function the prospect operates in, run the matching block below. Always quantify with numbers. Always end with the red-flag check (§16) before recommending automation.
Sales
Openers (ask all)
- Walk me through what happens between a lead landing and a closed deal.
- What's a typical week for your reps from Monday to Friday?
- If we deleted one task from every rep's day tomorrow, what should it be — and why?
- Where do deals usually slip or die?
- What's the single number — pipeline, win rate, cycle time, ramp — your VP gets graded on this quarter?
Conditional branches
- If they mention manual lead routing → Who does the routing? Rules today? How long does a lead sit before first touch?
- If they mention CRM hygiene → Which CRM? Last audit when? Which fields stalest? Who owns?
- If they mention slow follow-up → Current speed-to-lead? Target? What blocks reps from being faster?
- If they mention forecasting pain → How does CRO arrive at the number — spreadsheet, Clari, gut?
- If they mention notes don't make CRM → Call recording in use? Which tool? % of calls actually logged?
- If they mention proposal/quote delays → Walk through a quote from request to signed. Longest pause?
- If they mention outbound at volume → How personalized is each touch — templates, sequences, manual?
Quantify (force numbers)
- Leads/week? Avg response time vs goal?
- Calls per rep per day? How many logged?
- Time on CRM update per rep per day × reps × workdays.
- Win rate? Cycle time? ACV? Stage-to-stage slip rate?
Tool stack probes
- CRM — Salesforce, HubSpot, Pipedrive, Close, Attio, custom?
- Engagement — Outreach, Salesloft, Apollo, Clay?
- Call intel — Gong, Chorus, Fathom, Granola?
- Enrichment — Clearbit, Apollo, ZoomInfo, Cognism?
- CPQ — Salesforce CPQ, DealHub, PandaDoc, Google Docs?
Red flags (stop or scope down)
- Who owns CRM data quality? — if 'no one,' flag.
- Documented sales process? — if no, scope to one stage only.
- How many CRMs? — if >1, data consolidation pre-req.
- Anyone hired/fired around this in 6 mo? — political risk.
Customer Success
Openers (ask all)
- Walk me through a customer's first 90 days post-sale.
- How do you decide which customer to call today?
- What does a QBR look like end-to-end?
- How do you find out a customer is unhappy — before they tell you?
- What's your NRR target and the gap today?
Conditional branches
- If they mention onboarding pain → Written onboarding plan? Who tracks? Longest phase?
- If they mention renewals sneaking up → How far in advance does renewal motion start? Who triggers?
- If they mention health is gut feel → What signals — usage, support volume, exec engagement? Where do they live?
- If they mention QBR overhead → How long does QBR deck take? Who builds? % of QBRs that actually happen?
- If they mention churn surprises → Walk last 3 churns. Were there leading indicators in hindsight?
Quantify (force numbers)
- Accounts per CSM?
- QBR prep hrs/week × QBRs/qtr.
- Renewal trigger to close time?
- % accounts with updated success plan?
Tool stack probes
- CS platform — Gainsight, Catalyst, Vitally, ChurnZero, Planhat?
- Product analytics — Mixpanel, Amplitude, Heap, Pendo?
- Where does account data live — CRM, CSP, sheet?
Red flags (stop or scope down)
- Customer owner clear post-sale? — if not, flag.
- Renewal forecast exists? — if no, scope renewal visibility first.
- Product analytics wired up? — if no events, automation has nothing to read.
Customer Support
Openers (ask all)
- What's a typical day in your queue?
- What % of tickets are repeat questions?
- First-response and resolution time today?
- Where does your KB live and how stale is it?
- What gets escalated to eng/product, how often?
Conditional branches
- If they mention repeat questions → Top 5? KB article for each? Current?
- If they mention overnight pileup → 24/7 staffed? Overnight vs day volume?
- If they mention macro decay → Who owns macros? Update cadence?
- If they mention agent KB-search overhead → Time per ticket to find answer?
- If they mention escalation overflow → % tickets escalating? Tier-1 confidence threshold?
Quantify (force numbers)
- Tickets/day, /week?
- AHT vs target?
- Deflection rate (self-serve)?
- % of volume in top 10 categories?
Tool stack probes
- Helpdesk — Zendesk, Intercom, Front, Help Scout, Freshdesk?
- Existing AI — Fin, Zendesk AI, Forethought?
- KB — Notion, Confluence, helpdesk-native, scattered?
Red flags (stop or scope down)
- KB exists? — if no, build content first.
- Public or behind login? — affects RAG strategy.
- Regulated industry? — scope-down on what AI can answer.
Finance / Accounting / FP&A
Openers (ask all)
- Walk me through month-end close day 1 → final.
- Where do you lose most time — recon, AP, AR, reporting?
- Biggest manual lift between systems?
- What does your CFO ask for that always takes too long?
- Where do errors usually slip in?
Conditional branches
- If they mention AR/collections → Open invoices? DSO? Who follows up, how?
- If they mention AP/invoices → Where do invoices arrive? Who codes? 3-way match by hand?
- If they mention reconciliation → Which accounts? Auto vs manual? Exception handling?
- If they mention close speed → Days? Longest single task?
- If they mention reporting → Recurring report? Who builds, how long?
- If they mention expense mgmt → Card program? Manual reimbursements? Approval flow?
Quantify (force numbers)
- DSO, DPO, close days, # recons/month
- Hours/month on AR follow-ups
- % of invoices auto-matched
- # of JEs/close
Tool stack probes
- ERP/GL — QuickBooks, NetSuite, Sage Intacct, Xero, Dynamics?
- AP — Bill.com, Ramp Bill Pay, Tipalti, Stampli?
- AR — Upflow, Chaser, Versapay?
- Close — FloQast, Numeric, BlackLine?
- FP&A — Mosaic, Pigment, Cube, Adaptive?
Red flags (stop or scope down)
- Chart of accounts clean? — if messy, automation accelerates the mess.
- Books current? — if behind, fix first.
- Controller or just outsourced bookkeeper? — affects safe scope.
- Audited or SOX? — if yes, advisory + HITL only in critical controls.
Operations / RevOps / BizOps
Openers (ask all)
- Most painful handoff in the business right now?
- Worst spreadsheet — the one holding everything together?
- Two systems that should be talking but aren't?
- What's a 'status update' look like today?
- Monday morning report you wish you didn't build?
Conditional branches
- If they mention handoff failures → Between which teams? What gets dropped? How caught today?
- If they mention spreadsheet hell → How many people touch? How often breaks? Single source of truth elsewhere?
- If they mention tool sprawl → How many SaaS tools? Which integrate? Where copy/paste?
- If they mention manual reporting → Cadence? Audience? Decisions driven?
- If they mention undocumented process → Runbook anywhere? Or in heads?
Quantify (force numbers)
- Hours/week on handoff tracking, report assembly
- # tools in stack, integrations vs manual
- Fields copied per transaction × transactions/day
Tool stack probes
- iPaaS — Workato, Tray, Zapier, n8n, Make?
- Work mgmt — Asana, Linear, Monday, ClickUp, Jira?
- Data layer — warehouse? Reverse ETL? dbt?
- Reporting — Looker, Mode, Hex, Sheets?
Red flags (stop or scope down)
- Data warehouse exists? — if no & analytical need, scope it as pre-req.
- Ops owner? — if none, change-mgmt risk.
- Recent reorg in 90d? — scope down.
Marketing
Openers (ask all)
- How does a campaign go from idea to live?
- Where do creative requests get stuck?
- How do you attribute pipeline back to marketing?
- Recurring report you wish you didn't build?
- Content production pipeline?
Conditional branches
- If they mention creative intake chaos → Where requests come — form, Slack, email? Triage owner?
- If they mention attribution gaps → First-touch, multi-touch, none? Data home? Who runs it?
- If they mention content slowness → Typical lifecycle — brief→draft→review→publish?
- If they mention newsletter/lifecycle → Today's automation? Open/click rates? Personalization?
Quantify (force numbers)
- Requests/week, time-to-fulfill, time-to-publish
- % of pipeline attributed
- Content pieces/month, hours per piece
Tool stack probes
- Marketing automation — HubSpot, Marketo, Customer.io, Iterable, Braze?
- CMS — Webflow, WordPress, Sanity, Contentful?
- Attribution — GA4, Bizible, Dreamdata, HockeyStack?
Red flags (stop or scope down)
- Marketing & sales on same CRM? — if no, scope sync first.
- Brand guidelines documented? — if no, AI content goes off-rails.
HR / People / Recruiting
Openers (ask all)
- Noisiest part of your hiring funnel?
- Most repetitive HR ticket?
- Onboarding a new hire from offer → day 30?
- How do performance reviews actually run?
- Payroll / comp / benefits admin time per cycle?
Conditional branches
- If they mention recruiting black hole → Where do candidates wait longest — source-to-screen, screen-to-offer?
- If they mention ghosting candidates → Who owns comms? SLA?
- If they mention onboarding ad-hoc → Checklist exists? Owner? Tech vs people task split?
- If they mention review overhead → Cadence, inputs, hours per report?
- If they mention employee tickets → Inbound channels, volume, top categories?
Quantify (force numbers)
- Open reqs, time-to-fill, candidates per req
- Onboarding completion rate, time-to-productive
- HR ticket volume/month, top 5 categories
Tool stack probes
- ATS — Greenhouse, Ashby, Lever, Workable?
- HRIS — Rippling, Gusto, Deel, BambooHR, Workday, ADP?
- Performance — Lattice, 15Five, Culture Amp, Leapsome?
- Comp — Pave, Figures, Carta?
Red flags (stop or scope down)
- Job ladder/leveling documented? — if no, AI-assisted reviews premature.
- HR records consistent across systems? — if no, scope cleanup first.
- Sensitive data flows (comp/PHI/PII)? — extra scoping.
Legal / Compliance / GRC
Openers (ask all)
- Where does legal slow business most today?
- % of contracts on playbook vs custom redline?
- How do you handle vendor security questionnaires?
- Compliance evidence collection at audit time?
Conditional branches
- If they mention NDA/MSA delays → Avg turnaround? Stuck where? Self-serve possible?
- If they mention questionnaire load → How many/qtr? Who fills? Where answers live?
- If they mention audit prep → Continuous controls or fire drill? Evidence repo exists?
- If they mention playbook coverage → Documented? Maintainer?
Quantify (force numbers)
- Contracts/month, avg turnaround, % auto-executed
- Questionnaires/qtr, hours each
- Audit prep hrs/year
Tool stack probes
- CLM — Ironclad, LinkSquares, Lexion, ContractWorks, DocuSign CLM?
- Compliance — Drata, Vanta, Secureframe?
- Questionnaires — Loopio, Responsive, Conga?
Red flags (stop or scope down)
- GC in-house or fractional? — affects velocity.
- Regulated industry? — tight scope; advisory mode.
- Privilege concerns sending docs to LLM? — must scope private deploy.
Engineering / IT / DevOps
Openers (ask all)
- Most annoying IT/eng ticket that won't go away?
- Where does PR / code review get stuck?
- Time for a new engineer to ship their first PR?
- Incidents — runbook, tribal, both?
- IT request volume?
Conditional branches
- If they mention L1 ticket volume → Top 5 categories, % of total, avg resolution time
- If they mention PR slowness → Review turnaround, tooling, bottleneck people
- If they mention incident response → MTTR, runbook coverage, oncall rotation
- If they mention onboarding lag → Setup checklist? Internal docs? Codebase tour?
- If they mention shadow IT → SaaS visibility? Approvals?
Quantify (force numbers)
- Tickets/week, % auto-resolved
- PR cycle time, PRs/week
- MTTR, incidents/month
- Time-to-first-PR
Tool stack probes
- Source — GitHub, GitLab, Bitbucket?
- Observability — Datadog, NewRelic, Honeycomb, Grafana?
- Incident — PagerDuty, Rootly, FireHydrant, Incident.io?
- IT helpdesk — Jira SM, Atomicwork, Moveworks, Slack?
- SaaS mgmt — Zylo, Torii, BetterCloud?
Red flags (stop or scope down)
- Engineers willing to use AI agent? — culture check.
- Sensitive code / regulated data in repos? — what models can touch.
- Oncall burnout? — careful adding agent noise.
Exec / Owner-led SMB
Openers (ask all)
- Three things stealing the most hours from your week?
- What can only you do — and what shouldn't only you do?
- Walk me through a normal Monday.
- What do you wish you could see daily that you don't?
- Where do you trust your team, and where don't you?
Conditional branches
- If they mention inbox overload → Volume? % needing personal response? EA?
- If they mention calendar chaos → Who manages? Tools? Recurring meetings to cut?
- If they mention lack of visibility → What single metric would tell you Monday morning the business is fine?
- If they mention approval bottleneck → What approvals can be policy-based vs judgment?
- If they mention customer DMs → Volume, channels, triage possible?
Quantify (force numbers)
- Hrs/week on email, calendar, approvals, customer DMs, status meetings
- # recurring meetings, # approvals/week
Tool stack probes
- Email — Gmail, Outlook, Superhuman, Shortwave?
- Calendar — Reclaim, Motion, Clockwise, Cal.com?
- Comms — Slack, Telegram, iMessage, WhatsApp, SMS?
Red flags (stop or scope down)
- Anyone on the team side who could own the agent? — if no, system rots.
- SOPs documented? — if no, scope SOPs first.
- Prior automation attempt? — uncover failure pattern.
Automation readiness scoring rubric
decay: 12moScore each candidate process 1–5 on seven axes. Composite threshold determines audit recommendation.
Scoring rubric — to embed in the audit agent
For each candidate process, output a row with these seven scores (1–5 each, 5 best) plus rationale. Composite = sum, max 35. Recommend automation if composite ≥ 22AND no axis <2. Otherwise, route to anti-list or fix-process-first.
| Axis | What it measures | Signals that score 5 | Signals that score 1 |
|---|---|---|---|
| Frequency | How often the process runs | ≥ 5×/day per user; multiple users | ≤ 1×/month, one user |
| Time cost | Loaded $ / month spent on this manually | ≥ $2,000/mo recoverable | < $200/mo recoverable |
| Ambiguity | How well-specified the work is | Deterministic rules cover > 90% of cases | Each case is a judgment call |
| Stakes-of-error | Cost of a wrong action | Low (internal report, easily reversed) | High (legal, financial, customer-facing, irreversible) |
| Data availability | Required inputs already exist in systems | All inputs in APIs / structured | Inputs in heads or scattered docs |
| Tool-fit | Current stack supports automation cleanly | Native APIs, mature integration platform | Walled-garden tool, no API, no webhooks |
| Change-mgmt cost | Effort to get humans to adopt new flow | Sole owner who wants the help | Multi-team rollout, prior failed attempts, no champion |
Composite interpretation:
≥ 28 — recommend full agent build, rank top of report.
22–27 — recommend with caveats; spec HITL gates.
16–21 — recommend deterministic workflow (Zapier/n8n) instead of agent.
11–15 — fix-process-first; surface SOP or data hygiene need.
≤ 10 — leave human; place on anti-list with reasoning.
Decision tree — pick the right tool
decay: 12moGiven a scored process, decide between: full agent, deterministic workflow, RPA, fix-process-first, or leave-human.
Decision tree — system prompt block
- Is composite ≤ 10 or any axis = 1? → leave-human. Recommend it goes on the anti-list with reasoning.
- Is data availability ≤ 2? → fix-process-first. Recommend a data-consolidation / SOP-extraction project as Phase 1. Don't sell the agent yet.
- Is ambiguity ≥ 4 AND change-mgmt ≥ 4? → deterministic workflow (Zapier / n8n / Make). Lower ceiling on failure modes; faster adoption.
- Is the system fully closed (no API/webhook) but UI is stable? → browser-automation agent (Stagehand + Browserbase) only when there's no other path. Tag HITL.
- Is the task primarily document-in / structured-out (PDF / invoice / contract clause)? → Claude API + Skills (or Gumloop for batch). Default to deterministic for high-volume.
- Is the task conversational and customer-facing (support, outbound)? → buy the deflection layer (Intercom Fin / Zendesk AI / Retell). Don't internal-build.
- Is the task multi-step, multi-tool, with state across steps? → LangGraph + Temporal (or Mistral Workflows-pattern). Anchor durability before any agent logic.
- Is the task long-horizon (> tens of minutes per run)? → re-scope to smaller checkpoints. Long-horizon at scale is not reliable in 2026.
- Default fallback if ≥ 22 composite, no blockers: full agent with HITL on send/spend/legal/customer-visible writes.
Recommendation deliverable template
decay: 12moWhat an audit output looks like. Drop into the audit agent's report-generation prompt verbatim.
Recommendation template — emit one per top-5 process
# {{process_name}}
**Composite readiness:** {{composite}}/35 · **Confidence:** {{confidence}}%
**Owner / role:** {{owner_role}}
**Status:** {{status}}
**Reasoning:** {{one_sentence}}
## What you're doing today
{{manual_workflow_3_to_6_steps}}
## What you'll save
- Time: {{time_saved}} hrs/mo
- Money: ${{cost_saved}}/mo ({{cost_saved_assumption}})
- Annualized: ${{cost_saved_year}}
## How we'd build it
**Stack:** {{tools_needed}}
**Approach:** {{automation_approach}}
**Reliability today:** {{reliability_today}}/5 — {{reliability_rationale}}
**HITL:** {{hitl_mode}} — {{hitl_reason}}
**Estimated build:** {{build_hours}} hours · {{build_calendar_days}} calendar days
## Risks & failure modes
- {{primary_failure_mode}}
- {{mitigation}}
## What you'd own going forward
- {{maintenance_estimate}} per month · {{maintenance_owner}}
## Choose your path
- [ ] DIY — full build instructions inside
- [ ] We build for one shot: **${{one_shot_price}}**
- [ ] We build + maintain monthly: **${{monthly_price}}/mo** (recommended)Report header block (above all top-5 items)
# Audit for {{company_name}}, {{role}}
**Total recoverable:** ${{total_savings}}/mo · {{total_hours}} hrs/mo back
**Top 5 automations · {{anti_count}} traps avoided · {{wedge_count}} wedge plays**
Generated {{date}} · Confidence-weighted · See methodology at /researchRed flags during intake
decay: 12moSignals to scope-down or walk away. Walk-away signals override everything else.
Red flag detector — system prompt block
- "We don't have one source of truth for X." → scope-down to data consolidation Phase 1, or walk if they refuse.
- "The last team / agency got fired." → political risk. Get explicit on what failed, why, and who is the new sponsor. If you can't get a straight answer, walk.
- "We're in healthcare/store PHI / process payments / under SOX." → scope down. Advisory-only, or BAA-eligible infra (Azure OpenAI w/ BAA, AWS Bedrock, on-prem).
- "We need it deployed by Friday." → scope-down to one demo-able slice; set realistic 6–8 weeks; if they won't move, walk.
- "We tried [Zapier / a chatbot] and it failed." → spend 20 min on the post-mortem before any pitch.
- "Our team will adopt this — they're excited." → ask for proof (existing adoption rates, training plans, change champion). If absent, pilot with one team only.
- "Our process is too unique to automate." → probe for the one stable, repeated slice. Scope there or walk.
- "We don't want HITL — we want full automation." → educate on failure modes. If they refuse and stakes are high, walk.
- "Our IT team will help build it." → ask for named engineer with allocated hours. If vague, scope to end-to-end delivery.
- "Legal needs to approve everything before we move."→ map approval path, get named contact & turnaround estimate. Scope contract terms minimally.
- No clear process owner. → insist on a single named owner with decision authority before signing. If they can't name one, walk.
- No data inventory. → scope a paid 1–2 week discovery. Don't promise outcomes pre-mapping.
- No KPI for the process. → define one with them on the call. If they can't / won't commit, scope down or walk.
- Regulated industry + no compliance lead. → walk unless we have compliance expertise in-house.
- Multiple stakeholders openly disagree on the goal. → insist on a single decision-maker. Defer until they align internally.
- Recent reorg (< 90 days). → scope down to quick win that survives reorg risk, or wait.
- "We just need an MVP" with no defined outcome. → force a measurable target ("X goes from Y to Z in N weeks").
- "We're talking to 5 other agencies." → if they'll decide on lowest price, walk.
- "Process is in someone's head." → Phase 1 = SOP extraction. Charge for it. Or walk if they refuse.
- "We just got new leadership." → confirm new leader is sponsor. If not, wait.
- "We don't have budget approved." → real budget conversation before discovery. If they won't, walk.
- "Our customers are very different — every account is special." → scope to internal-ops automation, not customer-facing.
- "Can you guarantee X% accuracy?" → educate on evals, confidence thresholds, HITL. If they won't engage, walk.
- "Can it replace [person]?" → reframe to "augment + redirect." Be honest about what AI can/can't do for that role.
- "We want to own the IP / build in-house eventually." → price for transfer or be explicit on licensed vs transferred. Don't be surprised.
- Founder/CEO is on the call alone — no operator counterpart. → require an operator in discovery #2. If refused, scope tiny or walk.
- "We tried hiring and couldn't find anyone." → probe: is the role judgment-heavy (scope down) or repetitive (good fit)?
- "We don't have observability / logs / event tracking." → scope eval infra and instrumentation in Phase 1. Charge for it.
- Procurement-led conversation, no business owner present. → request a business stakeholder. If denied, walk.
- "We don't have time for discovery — just build it." → walk, or charge a non-refundable discovery fee that filters them out.
Pricing & business reality
decay: 6moWhat AI automation consultancies actually charge in 2025-2026. What converts. Cycle realities.
To populate from synthesis
The pricing-and-business research returns from a dedicated agent (raw-business-pricing.md). Synthesized into this section: shop benchmarks, engagement structures, hourly/per-build rates, sales cycle reality, free-audit conversion benchmarks, and the competitive map. Until that lands, the published number to anchor is Zapier Experts hourly $75–$250 zapier.com/experts, Make.com partner agency engagements $1k–$15k/build, and productized AI ops subscription $99–$2,000/mo range across emerging shops on r/AI_Agents.
What I can't verify yet · decay-watch
decay: 1moOpen questions. Honest holes. What gets stale fastest.
Holes in this draft
- Unit economics for the audit itself are modelled, not measured. Free-audit → paid conversion rate is a target, not a benchmark from a comparable shop. Need 100 audits to validate.
- Voice agent cost in production assumes Retell at $0.05/min — confirm at actual call durations (we expect drift to ~$0.08–$0.10/min with multimodal LLM costs included).
- Wedge-tier process count. 30+ is the target; the true number is whatever survives sales contact. Some never-automated entries will fail sales appeal; some will reveal they're rarely-automated by a niche vendor we missed.
- Pricing for "Pro Audit" ($499) is untested. The $99/mo subscription will validate first; the deeper paid audit comes second.
- Customer profile. "25–150 person services businesses" is the bet; validation requires 20+ paid customers before we know if it holds vs e-comm or solo SaaS.
Decay watch — re-run these first
- 1 month:capability matrix & vendor map (§6, §7). New SOTA every quarter; reliability scores drift.
- 3 months: failure-mode list (§9). Prompt-injection surface expands; mitigations mature.
- 6 months: process database (§3), wedge tier (§4), anti-list (§5), pricing-reality (§17). Wedges close when other vendors copy us.
- 12 months: thesis (§1), audit-job spec (§2), role pains (§10), discovery bank (§12), scoring rubric (§13), decision tree (§14), template (§15), red flags (§16). Stable, but check after first 100 paid engagements.
- Stable: operator ↔ vendor glossary (§11). Language drifts slowly; the underlying pains are decade-stable.
Source bibliography: /research/sources.md. Raw research files live in /research/raw-*.md for traceability.