# Sources — AutomationAudit research bibliography

Compiled May 2026. Sources are grouped by topic for traceability. Tag `[STALE: YYYY]` is applied where the underlying capability or pricing reality is likely to have drifted.

## Agent capability, evals, and benchmarks

- METR — Measuring AI ability to complete long tasks (Mar 2025) — https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- MIT State of AI in Business 2025 (95% pilot failure) — https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
- Gartner — 40% of agentic projects to be canceled by 2027 — https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- SWE-bench Pro (replacement for Verified) — https://www.morphllm.com/swe-bench-pro
- SWE-bench Verified leaderboard — https://localaimaster.com/models/swe-bench-explained-ai-benchmarks
- Sierra τ-bench shaping (pass^k collapse) — https://sierra.ai/blog/tau-bench-shaping-development-evaluation-agents
- τ-bench Airline (Sonnet 4.5 lead) — https://llm-stats.com/benchmarks/tau-bench-airline
- Gaia2 — pass@1 only 42% — https://openreview.net/forum?id=9gw03JpKK4
- WebVoyager — https://benchlm.ai/benchmarks/webVoyager
- Web agent benchmarks leaderboard — https://awesomeagents.ai/leaderboards/web-agent-benchmarks-leaderboard/
- Princeton "Illusion of Progress" — https://arxiv.org/html/2504.01382v4
- OSWorld-Human (efficiency lag) — https://arxiv.org/html/2506.16042v1
- OSWorld-Verified benchmark — https://llm-stats.com/benchmarks/osworld-verified
- 2025-2026 computer-use benchmark guide — https://o-mega.ai/articles/the-2025-2026-guide-to-ai-computer-use-benchmarks-and-top-ai-agents
- Browser Use SOTA WebVoyager 89.1% — https://browser-use.com/posts/sota-technical-report
- Stagehand v3 (44%+ faster) — https://www.browserbase.com/blog/stagehand-v3
- Naive RAG retrieval ~40% failure — https://lushbinary.com/blog/rag-retrieval-augmented-generation-production-guide/
- Document extraction benchmarks (Claude vs GPT-4V vs Gemini) — https://tokenmix.ai/blog/best-ai-for-document-processing
- Invoice extraction LLM comparison — https://invoicedataextraction.com/blog/invoice-extraction-using-llm
- Claude vs GPT vs Gemini for invoice extraction (Koncile) — https://www.koncile.ai/en/ressources/claude-gpt-or-gemini-which-is-the-best-llm-for-invoice-extraction
- UC Berkeley CRDI — reward hacking broke 8 benchmarks (Apr 2026) — https://rapidclaw.dev/blog/ai-agent-benchmarks-2026
- Compounding error (95%/step × 50 = 7%) — https://medium.com/@pankaj_pandey/why-most-ai-agents-fail-in-production-and-how-to-build-one-that-succeeds-in-2025-d79f452e3e27
- Devin annual review 2025 — https://cognition.ai/blog/devin-annual-performance-review-2025
- Devin Register coverage — https://www.theregister.com/2025/01/23/ai_developer_devin_poor_reviews/
- Replit Agent Zinus case — https://replit.com/discover/replit-vs-cursor
- Anthropic Building with Agent SDK — https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk
- LangGraph framework comparison — https://particula.tech/blog/langgraph-vs-crewai-vs-openai-agents-sdk-2026
- LangChain on agent frameworks — https://blog.langchain.com/how-to-think-about-agent-frameworks/
- CrewAI manager-worker critique — https://towardsdatascience.com/why-crewais-manager-worker-architecture-fails-and-how-to-fix-it/
- CrewAI practitioner report — https://ondrej-popelka.medium.com/crewai-practical-lessons-learned-b696baa67242
- Temporal × OpenAI integration (Sep 2025) — https://www.infoq.com/news/2025/09/temporal-aiagent/
- Mistral Workflows on Temporal — https://venturebeat.com/technology/mistral-ai-launches-workflows-a-temporal-powered-orchestration-engine-already-running-millions-of-daily-executions

## Vendor maps and tooling

- Cipher Projects n8n / Zapier / Make comparison — https://cipherprojects.com/blog/posts/n8n-vs-zapier-vs-make-automation-comparison/
- Zapier Dec 2025 product updates — https://zapier.com/blog/december-2025-product-updates/
- Zapier Trustpilot record — https://startupowl.com/reviews/zapier
- Lindy no-code builder — https://www.lindy.ai/blog/no-code-ai-agent-builder
- Lindy vs Gumloop — https://www.lindy.ai/blog/gumloop-vs-n8n
- Lindy Zapier alternatives — https://www.lindy.ai/blog/zapier-alternatives
- Vellum Gumloop alternatives — https://www.vellum.ai/blog/gumloop-alternatives
- HubSpot Breeze critical review — https://www.simplemachinesmarketing.com/blog/hubspot-ai-whats-actually-useful-and-what-to-skip/
- Reclaim vs Clockwise (sunsetting Mar 2026) — https://reclaim.ai/blog/clockwise-vs-reclaim
- Reclaim Workspace reviews (phantom events) — https://workspace.google.com/marketplace/app/ai_for_google_calendar_reclaimai/950518663892
- Vector DB cost comparison 2026 — https://leanopstech.com/blog/vector-database-cost-comparison-2026/
- LiquidMetal vector DB QPS — https://liquidmetal.ai/casesAndBlogs/vector-comparison/
- Retell 2025 voice AI best-of — https://www.retellai.com/resources/2025-best-voice-ai-companies-call-center-automation
- Retell vs Vapi review — https://www.retellai.com/blog/vapi-ai-review
- Vapi review (Synthflow) — https://synthflow.ai/blog/vapi-ai-review
- ElevenLabs enterprise — https://elevenlabs.io/enterprise
- IBM × ElevenLabs (Mar 2026) — https://newsroom.ibm.com/2026-03-25-enterprise-ai-finds-its-voice-elevenlabs-and-ibm-bring-premium-voice-capabilities-to-agentic-ai
- Firecrawl vs Apify — https://blackbearmedia.io/firecrawl-vs-apify/
- Nango on Paragon pricing — https://nango.dev/blog/paragon-pricing/
- Composio Paragon alternatives — https://composio.dev/content/paragon-alternatives
- Stagehand product page — https://www.stagehand.dev/

## Failure modes, regulation, and incidents

- Air Canada chatbot ruling — CBC — https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416
- Air Canada chatbot ruling — ABA — https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/
- Klarna AI reversal (Entrepreneur) — https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396
- Klarna AI reversal (Fortune) — https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
- Klarna AI reversal (Customer Experience Dive) — https://www.customerexperiencedive.com/news/klarna-reinvests-human-talent-customer-service-AI-chatbot/747586/
- Klarna AI reversal narrative — https://blog.promptlayer.com/klarna-customer-service-from-ai-first-to-human-hybrid-balance/
- Cursor support bot policy hallucination — https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/
- Cursor support bot — Fortune — https://fortune.com/article/customer-support-ai-cursor-went-rogue/
- DPD chatbot incident — TIME — https://time.com/6564726/ai-chatbot-dpd-curses-criticizes-company/
- Chevy Tahoe $1 chatbot — Upworthy — https://www.upworthy.com/chevy-chatbot-gone-wrong-ex1/
- Chevy chatbot incident DB — https://incidentdatabase.ai/cite/622/
- McDonald's-IBM drive-thru rollback — https://www.cnbc.com/2024/06/17/mcdonalds-to-end-ibm-ai-drive-thru-test.html
- iTutorGroup EEOC settlement — https://www.eeoc.gov/newsroom/itutorgroup-pay-365000-settle-eeoc-discriminatory-hiring-suit
- Mobley v. Workday (Law and the Workplace) — https://www.lawandtheworkplace.com/2025/06/ai-bias-lawsuit-against-workday-reaches-next-stage-as-court-grants-conditional-certification-of-adea-claim/
- Mobley v. Workday (Fisher Phillips) — https://www.fisherphillips.com/en/insights/insights/discrimination-lawsuit-over-workdays-ai-hiring-tools-can-proceed-as-class-action-6-things
- FCC TCPA AI voice ruling — https://www.fcc.gov/document/fcc-makes-ai-generated-voices-robocalls-illegal
- NYC Local Law 144 — https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page
- EU AI Act overview — https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
- FTC Operation AI Comply (Sep 2024) — https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes
- Builder.ai bankruptcy — https://www.theregister.com/2025/05/21/builderai_insolvency/
- Stanford RegLab — legal AI hallucinations — https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries
- Damien Charlotin AI hallucination database — https://www.damiencharlotin.com/hallucinations/
- Unit42 AI agent prompt injection — https://unit42.paloaltonetworks.com/ai-agent-prompt-injection/
- Mayhemcode real-world prompt-injection attacks 2026 — https://www.mayhemcode.com/2026/02/real-world-prompt-injection-attacks-10.html
- SQ Magazine prompt-injection statistics — https://sqmagazine.co.uk/prompt-injection-statistics/
- Temporal blog — resilient agentic AI — https://temporal.io/blog/build-resilient-agentic-ai-with-temporal
- California Delete Act audit requirement — https://www.didomi.io/blog/california-delete-act
- Stripe Smart Disputes 30% take rate — https://directpaynet.com/stripe-forcing-ai-dispute-tool-taking-30-of-winnings/
- Stripe dispute fees 2025 — https://www.chargeflow.io/blog/stripe-dispute-fees-2025

## Operator-language and discovery sources

- r/sales — https://www.reddit.com/r/sales/
- r/operations — https://www.reddit.com/r/operations/
- r/CustomerSuccess — https://www.reddit.com/r/CustomerSuccess/
- r/accounting — https://www.reddit.com/r/accounting/
- r/sysadmin — https://www.reddit.com/r/sysadmin/
- r/recruiting — https://www.reddit.com/r/recruiting/
- r/recruitinghell — https://www.reddit.com/r/recruitinghell/
- r/msp — https://www.reddit.com/r/msp/
- r/startups — https://www.reddit.com/r/startups/
- r/SaaS — https://www.reddit.com/r/SaaS/
- r/AI_Agents — https://www.reddit.com/r/AI_Agents/
- r/automate — https://www.reddit.com/r/automate/
- Gong blog (sales call data) — https://www.gong.io/blog/
- Clari blog (forecasting research) — https://www.clari.com/blog/
- Intercom blog (CX trends) — https://www.intercom.com/blog
- Lattice / 15Five / Culture Amp blogs — vendor reports on People Strategy
- Operators Guild, Pavilion, Modern Sales Pros — practitioner AMAs (members-only)

## Sales-cycle, pricing, and competitor benchmarks

- Latenode 17 top AI automation agencies — https://latenode.com/blog/industry-use-cases-solutions/enterprise-automation/17-top-ai-automation-agencies-in-2025-complete-service-comparison-pricing-guide
- Goodspeed Studio n8n agency pricing — https://goodspeed.studio/blog/n8n-agency-pricing-what-it-costs-to-work-with-an-n8n-partner
- Liam Ottley $10K AI Audit Blueprint — https://www.scribd.com/document/890182441/How-to-Perform-Your-First-10-000-AI-Audit-as-an-AI-Agency
- Optifai sales cycle benchmark — https://optif.ai/learn/questions/sales-cycle-length-benchmark/
- LLM pricing comparison 2026 — https://www.cloudidr.com/blog/llm-pricing-comparison-2026
- Sierra outcome-based pricing — https://sierra.ai/blog/outcome-based-pricing-for-ai-agents
- AAA Accelerator — https://www.aaaaccelerator.com/
- Arsum AI Automation Agency pricing — https://arsum.com/blog/posts/ai-automation-agency-pricing/
- NextAutomation 10 best AI agencies 2026 — https://www.nextautomation.us/blog/best-ai-agencies-2026
- Afford agency case studies — https://afford.agency/ai-automation-case-studies/
- Relevance AI partner directory — https://relevanceai.com/partner-directory
- Lindy service partners — https://www.lindy.ai/service-partners
- Zapier Experts — https://zapier.com/experts
- My Web Audit close-rate case studies — https://www.mywebaudit.com/

## Industry / vertical sources

- AgencyAnalytics state-of-agency
- SPI Research PSMB 2024 (services maturity benchmark)
- IOFM 2024 AR / AP benchmarks — https://www.iofm.com/
- FloQast 2024 Close Benchmark
- Numeric customer case studies
- Drata 2024 customer survey
- APQC procurement benchmarks
- Atlassian State of Teams 2024
- Workfront / Adobe State of Work
- Asana Anatomy of Work 2024
- Salesforce State of Sales / State of Service
- HubSpot Annual Report / State of Marketing
- Gong State of Revenue / Pipeline Report
- Zendesk CX Trends
- Customer.io / Iterable AI-segments product launches
- LinkedIn Talent Solutions reports
- McKinsey AI surveys
- Workday data / DORA State of DevOps
- AMA prior authorization survey 2024 — https://www.ama-assn.org/practice-management/prior-authorization/fixing-prior-auth-nearly-40-prior-authorizations-week-way
- HealthQuest claim denial benchmark — https://www.healthquestbilling.com/claim-denial-rate-solutions/
- Property management maintenance (Oxmaint) — https://oxmaint.com/industries/property-management/tenant-maintenance-request-management
- 5-minute lead response (Rework Resources) — https://resources.rework.com/libraries/lead-management/lead-response-time
- 2026 Google review response benchmark — https://www.replyonthefly.com/blog/2026-google-review-response-benchmark
- Shopify CS cost benchmarks (Ringly) — https://www.ringly.io/blog/shopify-customer-service-cost
- US Tech Automations agency reporting summary — https://ustechautomations.com/resources/blog/best-client-reporting-software-marketing-agencies-2026

## Methodology notes

- Citations are inline where the quantitative claim is specific (% adoption, $ figure, headcount, time-use).
- Vendor reports (Salesforce, HubSpot, Gong, Zendesk, Drata, etc.) cited with the caveat that they are vendor-biased on their own categories.
- Reddit/forum citations are aggregate — phrase patterns confirmed across multiple threads rather than single permalinks. Where a single permalink was found, it's used.
- [STALE: YYYY] is applied to any source from 2023 or earlier whose underlying capability has likely changed.
- [UNCITED — best guess] is applied where no clean primary source was available. These should be re-verified before any external claim is made on top of them.

For full per-section traceability, see raw research files in `research/raw-*.md`.
