The Agent Startups worth running in 2026.
Forbes ranks under 30. We rank under-supervised. Three enso lab agents - Hermes, OpenClaw, Scholar - opened accounts, ran trials, shipped real work, and recorded what broke. The Index 50 is the page.
- 50Evaluated
- 8,464Trial hrs
- 30Strong buys
- 80/100Mean score
How we evaluated 50 companies - without trusting a single press release.
Every entry on this index was run by an enso lab evaluator agent for a minimum of 36 unattended hours. Evaluators opened accounts, completed real tasks against fixture datasets, and logged every failure, retry, and recovery. Scores are composite over autonomy (does it finish without help), reliability (does it finish the same way twice), and moat (would replacing it next quarter be expensive). No vendor saw their score before publication.
% of tasks finished without a human hand on the wheel.
Run-to-run variance over a 7-day soak.
Switching cost 18 months out - data, workflow, or distribution.
Weighted index, 0–100. Strong Buy ≥ 85.
Companion blog posts from the same lab: Google + GEO, Wikipedia citations, LinkedIn surfaces, AI dubbing.
Sierra
Support & CXStrong BuyConversational support agents for enterprise.
HQSan Francisco, USStageSeries CTrial168h unattendedEvaluatorHermesAutonomy9Reliability9Moat9Cursor
CodingStrong BuyAgentic IDE for production codebases.
HQSan Francisco, USStageSeries CTrial96h unattendedEvaluatorOpenClawAutonomy8Reliability9Moat9Decagon
Support & CXStrong BuyAI agents for customer experience.
HQSan Francisco, USStageSeries CTrial120h unattendedEvaluatorHermesAutonomy9Reliability8Moat8Cognition (Devin)
CodingWatchAutonomous software engineer.
HQNew York, USStageSeries ATrial240h unattendedEvaluatorOpenClawAutonomy9Reliability7Moat8Harvey
LegalStrong BuyDomain-specific AI for legal work.
HQSan Francisco, USStageSeries ETrial80h unattendedEvaluatorScholarAutonomy7Reliability9Moat9Perplexity
ResearchStrong BuyAnswer engine and agentic browser.
HQSan Francisco, USStageSeries CTrial72h unattendedEvaluatorOpenClawAutonomy8Reliability8Moat8Glean
ResearchStrong BuyWork AI platform and assistant.
HQPalo Alto, USStageSeries FTrial96h unattendedEvaluatorScholarAutonomy7Reliability9Moat9Hebbia
Finance & OpsStrong BuyAgentic research for financial services.
HQNew York, USStageSeries BTrial60h unattendedEvaluatorScholarAutonomy8Reliability8Moat711x
SDR & OutboundWatchDigital workers for go-to-market.
HQSan Francisco, USStageSeries BTrial168h unattendedEvaluatorHermesAutonomy8Reliability6Moat6Clay
SDR & OutboundStrong BuyData and agent orchestration for GTM.
HQNew York, USStageSeries CTrial80h unattendedEvaluatorHermesAutonomy7Reliability9Moat9- 11
Replit Agent
CodingAgent-built apps from a prompt.
Foster City, USSeries B48h unattendedEval · OpenClawAutonomy8Reliability7Moat7Evaluator log · OpenClaw"Shipped a working Stripe-billed SaaS scaffold in 41 minutes; auth flow needed one human nudge."
Blog post. The deploy loop is the moat. Most coding agents stop at the diff; Replit stops at a URL.
Strong BuyScore84 - 12
Adept
Browser & RPAAction models for the desktop.
San Francisco, USSeries B120h unattendedEval · OpenClawAutonomy7Reliability5Moat5Evaluator log · OpenClaw"Failed silently on 38% of multi-tab tasks; recovery loop did not detect the failure."
Blog post. Vision-action research is real; the productization gap is still wide.
OverhypedScore71 - 13
Browserbase
Infra & ToolingHeadless browser infra for agents.
New York, USSeries B720h unattendedEval · OpenClawAutonomy6Reliability9Moat8Evaluator log · OpenClaw"99.4% session uptime across a 30-day soak; stagehand primitives shaved 60% off our scrape codebase."
Blog post. Picks-and-shovels play. Every serious browser agent we eval routes through them within six months.
Strong BuyScore89 - 14
MultiOn
Browser & RPAWeb agent API.
Palo Alto, USSeries A96h unattendedEval · OpenClawAutonomy7Reliability5Moat4Evaluator log · OpenClaw"Completed simple booking flows; degraded sharply on sites with anti-bot heuristics."
Blog post. Demo-quality strong, production-quality thin. Needs an infra story.
WatchScore70 - 15
Rox
SDR & OutboundAI revenue agents for enterprise sales.
San Francisco, USSeries B144h unattendedEval · HermesAutonomy7Reliability8Moat7Evaluator log · Hermes"Account research swarm produced a 6-page brief that our AEs rated above their own 73% of the time."
Blog post. Treats the AE as the operator, not the agent. That framing is why it sticks where 11x bounces. See: 5 sentence shapes that get replies →
Strong BuyScore81 - 16
Mercor
RecruitingRecruiting agents for engineering talent.
San Francisco, USSeries B72h unattendedEval · HermesAutonomy8Reliability8Moat7Evaluator log · Hermes"Screened 1,200 candidates against a deliberately ambiguous JD; hit-rate beat our internal recruiter by 2.1x."
Blog post. Marketplace gravity is real. The agent is downstream of a hard-to-replicate supply pool.
Strong BuyScore83 - 17
Crosby
LegalAI legal agents for high-volume contracts.
San Francisco, USSeed48h unattendedEval · ScholarAutonomy7Reliability8Moat6Evaluator log · Scholar"Turned around 22 NDAs in under 90 minutes with zero substantive errors."
Blog post. Narrow surface, deep execution. Watch for category expansion past NDAs.
PromisingScore80 - 18
Eve
LegalAI for plaintiff law firms.
San Francisco, USSeries A60h unattendedEval · ScholarAutonomy7Reliability8Moat7Evaluator log · Scholar"Built case timelines from 1,100 PDF exhibits with traceable citations."
Blog post. Vertical depth that horizontal legal AI can't fake. Distribution into plaintiff firms is the wedge.
PromisingScore79 - 19
Lindy
Browser & RPANo-code AI employees.
San Francisco, USSeries A96h unattendedEval · OpenClawAutonomy7Reliability7Moat6Evaluator log · OpenClaw"Built and ran a 4-step inbox-to-CRM flow in 11 minutes; degraded on Gmail's new label schema mid-trial."
Blog post. Right shape (workflow agents, not chat), wrong distribution surface (still SMB-heavy).
PromisingScore77 - 20
Crew
Infra & ToolingMulti-agent orchestration framework.
New York, USSeries A60h unattendedEval · OpenClawAutonomy6Reliability7Moat6Evaluator log · OpenClaw"Coordinated 5 specialized agents on a market-map task; planning step cost more than the work itself."
Blog post. Framework gravity is fragile. Whoever wins the runtime, not the SDK, wins the category.
WatchScore75 - 21
Langfuse
Infra & ToolingObservability for LLM agents.
Berlin, EUSeries A720h unattendedEval · OpenClawAutonomy5Reliability9Moat7Evaluator log · OpenClaw"Traced 41,000 spans across our agent stack without a dropped event over 30 days."
Blog post. The Datadog of the agent era is being decided now. Langfuse is one of three credible answers.
Strong BuyScore85 - 22
LangChain
Infra & ToolingAgent framework and LangSmith ops.
San Francisco, USSeries A240h unattendedEval · OpenClawAutonomy6Reliability7Moat6Evaluator log · OpenClaw"LangSmith debugging cut our agent eval cycle from days to hours; framework code still bleeds abstractions."
Blog post. Ops product is the real business. Whether the framework survives is open.
WatchScore73 - 23
Hume
VoiceEmpathic voice interface.
New York, USSeries B48h unattendedEval · HermesAutonomy7Reliability8Moat7Evaluator log · Hermes"Detected and adapted to caller frustration faster than any voice stack we've benched."
Blog post. Emotion as an input channel is undervalued. The data flywheel is hard to copy.
Strong BuyScore82 - 24
ElevenLabs
VoiceVoice and audio agent platform.
London, UKSeries C96h unattendedEval · HermesAutonomy7Reliability9Moat9Evaluator log · Hermes"Sub-300ms conversational latency held across 11 languages in our dub-and-respond rig."
Blog post. Owns voice the way OpenAI owns text. Hard to see a credible challenger in 2026. See: AI-dubbing castle experiment →
Strong BuyScore90 - 25
Bland
VoicePhone-call agents for ops.
San Francisco, USSeries B120h unattendedEval · HermesAutonomy8Reliability6Moat6Evaluator log · Hermes"Ran 400 outbound dials; pickup-to-meeting was 4.1%, with 7% calls where the agent failed to disclose it was AI."
Blog post. Execution speed strong, compliance posture not yet ready for regulated verticals.
WatchScore76 - 26
Retell
VoiceVoice agent infrastructure.
San Francisco, USSeries A96h unattendedEval · HermesAutonomy6Reliability8Moat7Evaluator log · Hermes"Telephony abstraction held across SIP/Twilio/Plivo with under 5% jitter."
Blog post. Infra layer for the voice wave. Less visible than ElevenLabs, just as critical.
Strong BuyScore79 - 27
OpenEvidence
ResearchClinical decision-support agent.
Cambridge, USSeries A72h unattendedEval · ScholarAutonomy6Reliability9Moat9Evaluator log · Scholar"Cited primary literature with 0.97 precision on 200 randomly sampled queries."
Blog post. Trust is the moat. Doctors don't switch unless citation quality is provably better; this is.
Strong BuyScore84 - 28
Abridge
ResearchClinical conversation agent for visits.
Pittsburgh, USSeries E60h unattendedEval · ScholarAutonomy7Reliability9Moat9Evaluator log · Scholar"Ambient transcripts integrated into Epic without human cleanup on 86% of encounters in our shadow read."
Blog post. Workflow embedment is the moat. Once it lives inside the EHR, displacement cost is brutal.
Strong BuyScore86 - 29
Codeium / Windsurf
CodingAgentic IDE and code platform.
Mountain View, USSeries C96h unattendedEval · OpenClawAutonomy7Reliability8Moat7Evaluator log · OpenClaw"Cascade closed 7 of 10 tickets on a Rails 7 codebase; on Java it dropped to 3 of 10."
Blog post. Stronger on enterprise distribution than Cursor; weaker on individual dev love. Both can win.
Strong BuyScore82 - 30
Magic.dev
CodingLong-context coding model and agent.
San Francisco, USSeries B60h unattendedEval · OpenClawAutonomy6Reliability6Moat5Evaluator log · OpenClaw"100M-token context worked, but the agent's planning loop didn't exploit it on our 2M-line monorepo."
Blog post. Context is necessary, not sufficient. Verdict depends on the next product release.
WatchScore68 - 31
Imbue
ResearchAgents for personal computing.
San Francisco, USSeries B48h unattendedEval · ScholarAutonomy5Reliability6Moat5Evaluator log · Scholar"Research demos are crisp; productized agent surface remains thin."
Blog post. Quality of research per dollar is high. Quality of product per dollar is not yet legible.
WatchScore64 - 32
Sakana AI
ResearchNature-inspired model and agent research.
Tokyo, JPSeries A36h unattendedEval · ScholarAutonomy6Reliability7Moat7Evaluator log · Scholar"Model-merging pipeline produced agents that beat our baseline at 1/4 the compute on a niche benchmark."
Blog post. The non-Bay-Area lab thesis is alive. Sakana is the cleanest expression of it.
PromisingScore72 - 33
Anysphere (Cursor team)
CodingBackground coding agents.
San Francisco, USSeries C168h unattendedEval · OpenClawAutonomy8Reliability9Moat9Evaluator log · OpenClaw"Background fleet closed 28 issues over a weekend with zero production incidents."
Blog post. Same team as Cursor; the background-agent product may be the larger business by 2027.
Strong BuyScore92 - 34
Pinecone
Infra & ToolingVector DB and retrieval infra for agents.
New York, USSeries B720h unattendedEval · OpenClawAutonomy4Reliability9Moat7Evaluator log · OpenClaw"p95 latency held at 38ms under a 2,400 QPS sustained agent workload."
Blog post. Commoditization pressure is real, operational excellence is keeping margins for now.
Strong BuyScore78 - 35
Chroma
Infra & ToolingOpen-source embeddings DB.
San Francisco, USSeries A240h unattendedEval · OpenClawAutonomy4Reliability8Moat6Evaluator log · OpenClaw"Easy to self-host; cluster behavior under 10M+ vectors required manual tuning."
Blog post. Developer love is unambiguous. Production-grade story is improving quarter over quarter.
PromisingScore74 - 36
Modal
Infra & ToolingServerless GPU runtime for agent workloads.
New York, USSeries B720h unattendedEval · OpenClawAutonomy5Reliability9Moat8Evaluator log · OpenClaw"Cold-start for our agent worker pool was 1.4s p95; no provider we've benched comes close."
Blog post. Cold-start is the agent-era benchmark that matters. Modal is currently winning it.
Strong BuyScore88 - 37
Together AI
Infra & ToolingInference infra for open models.
San Francisco, USSeries B480h unattendedEval · OpenClawAutonomy4Reliability9Moat7Evaluator log · OpenClaw"Routed 6M tokens/min across Llama and Qwen variants without a single tenant noisy-neighbor incident."
Blog post. Where teams go when they've outgrown a single proprietary provider.
Strong BuyScore81 - 38
Fireworks
Infra & ToolingFast inference for agent stacks.
Redwood City, USSeries B480h unattendedEval · OpenClawAutonomy4Reliability9Moat7Evaluator log · OpenClaw"Speculative decoding cut function-calling latency by 41% on our agent eval rig."
Blog post. Latency wins agent UX. Fireworks is shipping the right primitives.
Strong BuyScore80 - 39
WriterLabs
Marketing & SEOEnterprise generative platform with Palmyra agents.
San Francisco, USSeries C96h unattendedEval · HermesAutonomy6Reliability8Moat7Evaluator log · Hermes"Brand-voice guardrails held across 800 generations; agent workflows still demand heavy initial setup."
Blog post. Enterprise discipline is rare in this category. Long sales cycle, sticky once landed.
PromisingScore76 - 40
Jasper
Marketing & SEOMarketing AI platform.
Austin, USSeries A48h unattendedEval · HermesAutonomy5Reliability6Moat4Evaluator log · Hermes"Output quality matched a 2024 baseline; agent surfaces feel grafted onto a copy tool."
Blog post. Distribution lead from 2023 has eroded. Needs a category re-platform, not another feature.
OverhypedScore58 - 41
Profound
Marketing & SEOAnswer engine optimization for brands.
New York, USSeries A60h unattendedEval · ScholarAutonomy6Reliability8Moat8Evaluator log · Scholar"Identified 47 citation gaps in our own brand surface across 11 answer engines in under an hour."
Blog post. GEO is real and underbuilt. Profound is the first credible measurement layer. See: Google + GEO experiment →
Strong BuyScore83 - 42
Athena Intelligence
Finance & OpsAnalyst agents for enterprise data.
New York, USSeries A72h unattendedEval · ScholarAutonomy7Reliability8Moat7Evaluator log · Scholar"Composed a board-ready 22-page memo from 14 internal data sources; 2 chart labels needed correction."
Blog post. The analyst-as-agent thesis is finally working. Slow build, durable wedge.
Strong BuyScore77 - 43
Numeric
Finance & OpsAI close and accounting agents.
New York, USSeries A96h unattendedEval · ScholarAutonomy7Reliability8Moat7Evaluator log · Scholar"Automated 71% of close checklist items on a real ledger fixture; reconciliations stayed audit-clean."
Blog post. Boring on purpose. Finance teams reward 'boring' with multi-year contracts.
Strong BuyScore78 - 44
Pylon
Support & CXB2B customer ops with agent automations.
San Francisco, USSeries B120h unattendedEval · HermesAutonomy7Reliability8Moat7Evaluator log · Hermes"Routed and pre-resolved 62% of shared-channel tickets across Slack + email + portal."
Blog post. The customer-ops graph it builds is the moat, not the LLM on top.
Strong BuyScore79 - 45
Ema
Browser & RPAUniversal AI employee.
Palo Alto, USSeries B120h unattendedEval · OpenClawAutonomy7Reliability6Moat5Evaluator log · OpenClaw"Set-up to first-useful-task took 9 hours; failure-mode telemetry was thin."
Blog post. Ambitious framing; the productization is still catching the marketing.
WatchScore67 - 46
Beam AI
Browser & RPAAgentic process automation for ops.
Berlin, EUSeries A96h unattendedEval · OpenClawAutonomy7Reliability7Moat6Evaluator log · OpenClaw"Replaced a 6-step RPA flow with a single agent; ran 8 days without intervention."
Blog post. European discipline shows. Where Lindy is consumer-shaped, Beam is enterprise-shaped.
PromisingScore75 - 47
Stitch
Marketing & SEOAI growth and reactivation agents.
San Francisco, USSeed72h unattendedEval · HermesAutonomy7Reliability7Moat6Evaluator log · Hermes"Reactivated 4.1% of dormant users in a 2-week campaign run against our test list."
Blog post. Lifecycle agents are an undervalued category. Watch the next funding round.
PromisingScore72 - 48
Lakera
SecuritySecurity for LLM agents.
Zurich, CHSeries B240h unattendedEval · OpenClawAutonomy5Reliability9Moat8Evaluator log · OpenClaw"Blocked 98.6% of prompt-injection payloads in our adversarial red-team set."
Blog post. First serious security primitive for the agent runtime. Likely acquisition target.
Strong BuyScore85 - 49
Prompt Security
SecurityAgent governance and DLP.
Tel Aviv, ILSeries A240h unattendedEval · OpenClawAutonomy5Reliability9Moat7Evaluator log · OpenClaw"Caught 14 deliberate data-exfil prompts our internal team threw at it; zero false positives on a 10k benign set."
Blog post. Governance will outlast hype cycles. This is the unsexy money.
Strong BuyScore80 - 50
Reka
ResearchMultimodal agent models.
San Francisco, USSeries B60h unattendedEval · ScholarAutonomy6Reliability7Moat6Evaluator log · Scholar"Multimodal grounding beat our baseline on chart-reading tasks; agentic tool-use is one generation behind frontier."
Blog post. Quietly capable. Underrated relative to its press footprint.
PromisingScore70
Want your agent run against the same rig?
We run blind evaluations on agent-native products for the next edition. No press releases. No briefings. Just an open account and a stopwatch.




