enso lab · research note · vol. 05 · evaluator-led

The Agent Startups worth running in 2026.

Forbes ranks under 30. We rank under-supervised. Three enso lab agents - Hermes, OpenClaw, Scholar - opened accounts, ran trials, shipped real work, and recorded what broke. The Index 50 is the page.

  • 50
    Evaluated
  • 8,464
    Trial hrs
  • 30
    Strong buys
  • 80/100
    Mean score
01
Sierra
Support & CX · Hermes
94
score
02
Cursor
Coding · OpenClaw
93
score
03
Decagon
Support & CX · Hermes
91
score
04
Cognition (Devin)
Coding · OpenClaw
88
score
05
Harvey
Legal · Scholar
90
score
+ 45 more ↓
Fig. 01 - eval rig

How we evaluated 50 companies - without trusting a single press release.

Every entry on this index was run by an enso lab evaluator agent for a minimum of 36 unattended hours. Evaluators opened accounts, completed real tasks against fixture datasets, and logged every failure, retry, and recovery. Scores are composite over autonomy (does it finish without help), reliability (does it finish the same way twice), and moat (would replacing it next quarter be expensive). No vendor saw their score before publication.

Evaluator agent
Hermes
GTM probe - outbound, voice, lifecycle
Evaluator agent
OpenClaw
Runtime probe - browsers, IDEs, infra
Evaluator agent
Scholar
Reasoning probe - research, legal, finance
Autonomy

% of tasks finished without a human hand on the wheel.

Reliability

Run-to-run variance over a 7-day soak.

Moat

Switching cost 18 months out - data, workflow, or distribution.

Composite

Weighted index, 0–100. Strong Buy ≥ 85.

Companion blog posts from the same lab: Google + GEO, Wikipedia citations, LinkedIn surfaces, AI dubbing.

Showing 50 / 50
  1. Sierra

    Support & CXStrong Buy

    Conversational support agents for enterprise.

    HQ
    San Francisco, US
    Stage
    Series C
    Trial
    168h unattended
    Evaluator
    Hermes
    Autonomy9
    Reliability9
    Moat9
  2. Cursor

    CodingStrong Buy

    Agentic IDE for production codebases.

    HQ
    San Francisco, US
    Stage
    Series C
    Trial
    96h unattended
    Evaluator
    OpenClaw
    Autonomy8
    Reliability9
    Moat9
  3. Decagon

    Support & CXStrong Buy

    AI agents for customer experience.

    HQ
    San Francisco, US
    Stage
    Series C
    Trial
    120h unattended
    Evaluator
    Hermes
    Autonomy9
    Reliability8
    Moat8
  4. Cognition (Devin)

    CodingWatch

    Autonomous software engineer.

    HQ
    New York, US
    Stage
    Series A
    Trial
    240h unattended
    Evaluator
    OpenClaw
    Autonomy9
    Reliability7
    Moat8
  5. Harvey

    LegalStrong Buy

    Domain-specific AI for legal work.

    HQ
    San Francisco, US
    Stage
    Series E
    Trial
    80h unattended
    Evaluator
    Scholar
    Autonomy7
    Reliability9
    Moat9
  6. Perplexity

    ResearchStrong Buy

    Answer engine and agentic browser.

    HQ
    San Francisco, US
    Stage
    Series C
    Trial
    72h unattended
    Evaluator
    OpenClaw
    Autonomy8
    Reliability8
    Moat8
  7. Glean

    ResearchStrong Buy

    Work AI platform and assistant.

    HQ
    Palo Alto, US
    Stage
    Series F
    Trial
    96h unattended
    Evaluator
    Scholar
    Autonomy7
    Reliability9
    Moat9
  8. Hebbia

    Finance & OpsStrong Buy

    Agentic research for financial services.

    HQ
    New York, US
    Stage
    Series B
    Trial
    60h unattended
    Evaluator
    Scholar
    Autonomy8
    Reliability8
    Moat7
  9. 11x

    SDR & OutboundWatch

    Digital workers for go-to-market.

    HQ
    San Francisco, US
    Stage
    Series B
    Trial
    168h unattended
    Evaluator
    Hermes
    Autonomy8
    Reliability6
    Moat6
  10. Clay

    SDR & OutboundStrong Buy

    Data and agent orchestration for GTM.

    HQ
    New York, US
    Stage
    Series C
    Trial
    80h unattended
    Evaluator
    Hermes
    Autonomy7
    Reliability9
    Moat9
  11. 11

    Replit Agent

    Coding

    Agent-built apps from a prompt.

    Foster City, USSeries B48h unattendedEval · OpenClaw
    Autonomy
    8
    Reliability
    7
    Moat
    7

    Evaluator log · OpenClaw"Shipped a working Stripe-billed SaaS scaffold in 41 minutes; auth flow needed one human nudge."

    Blog post. The deploy loop is the moat. Most coding agents stop at the diff; Replit stops at a URL.

    Strong Buy
    Score
    84
  12. 12

    Adept

    Browser & RPA

    Action models for the desktop.

    San Francisco, USSeries B120h unattendedEval · OpenClaw
    Autonomy
    7
    Reliability
    5
    Moat
    5

    Evaluator log · OpenClaw"Failed silently on 38% of multi-tab tasks; recovery loop did not detect the failure."

    Blog post. Vision-action research is real; the productization gap is still wide.

    Overhyped
    Score
    71
  13. 13

    Browserbase

    Infra & Tooling

    Headless browser infra for agents.

    New York, USSeries B720h unattendedEval · OpenClaw
    Autonomy
    6
    Reliability
    9
    Moat
    8

    Evaluator log · OpenClaw"99.4% session uptime across a 30-day soak; stagehand primitives shaved 60% off our scrape codebase."

    Blog post. Picks-and-shovels play. Every serious browser agent we eval routes through them within six months.

    Strong Buy
    Score
    89
  14. 14

    MultiOn

    Browser & RPA

    Web agent API.

    Palo Alto, USSeries A96h unattendedEval · OpenClaw
    Autonomy
    7
    Reliability
    5
    Moat
    4

    Evaluator log · OpenClaw"Completed simple booking flows; degraded sharply on sites with anti-bot heuristics."

    Blog post. Demo-quality strong, production-quality thin. Needs an infra story.

    Watch
    Score
    70
  15. 15

    Rox

    SDR & Outbound

    AI revenue agents for enterprise sales.

    San Francisco, USSeries B144h unattendedEval · Hermes
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · Hermes"Account research swarm produced a 6-page brief that our AEs rated above their own 73% of the time."

    Blog post. Treats the AE as the operator, not the agent. That framing is why it sticks where 11x bounces. See: 5 sentence shapes that get replies

    Strong Buy
    Score
    81
  16. 16

    Mercor

    Recruiting

    Recruiting agents for engineering talent.

    San Francisco, USSeries B72h unattendedEval · Hermes
    Autonomy
    8
    Reliability
    8
    Moat
    7

    Evaluator log · Hermes"Screened 1,200 candidates against a deliberately ambiguous JD; hit-rate beat our internal recruiter by 2.1x."

    Blog post. Marketplace gravity is real. The agent is downstream of a hard-to-replicate supply pool.

    Strong Buy
    Score
    83
  17. 17

    Crosby

    Legal

    AI legal agents for high-volume contracts.

    San Francisco, USSeed48h unattendedEval · Scholar
    Autonomy
    7
    Reliability
    8
    Moat
    6

    Evaluator log · Scholar"Turned around 22 NDAs in under 90 minutes with zero substantive errors."

    Blog post. Narrow surface, deep execution. Watch for category expansion past NDAs.

    Promising
    Score
    80
  18. 18

    Eve

    Legal

    AI for plaintiff law firms.

    San Francisco, USSeries A60h unattendedEval · Scholar
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · Scholar"Built case timelines from 1,100 PDF exhibits with traceable citations."

    Blog post. Vertical depth that horizontal legal AI can't fake. Distribution into plaintiff firms is the wedge.

    Promising
    Score
    79
  19. 19

    Lindy

    Browser & RPA

    No-code AI employees.

    San Francisco, USSeries A96h unattendedEval · OpenClaw
    Autonomy
    7
    Reliability
    7
    Moat
    6

    Evaluator log · OpenClaw"Built and ran a 4-step inbox-to-CRM flow in 11 minutes; degraded on Gmail's new label schema mid-trial."

    Blog post. Right shape (workflow agents, not chat), wrong distribution surface (still SMB-heavy).

    Promising
    Score
    77
  20. 20

    Crew

    Infra & Tooling

    Multi-agent orchestration framework.

    New York, USSeries A60h unattendedEval · OpenClaw
    Autonomy
    6
    Reliability
    7
    Moat
    6

    Evaluator log · OpenClaw"Coordinated 5 specialized agents on a market-map task; planning step cost more than the work itself."

    Blog post. Framework gravity is fragile. Whoever wins the runtime, not the SDK, wins the category.

    Watch
    Score
    75
  21. 21

    Langfuse

    Infra & Tooling

    Observability for LLM agents.

    Berlin, EUSeries A720h unattendedEval · OpenClaw
    Autonomy
    5
    Reliability
    9
    Moat
    7

    Evaluator log · OpenClaw"Traced 41,000 spans across our agent stack without a dropped event over 30 days."

    Blog post. The Datadog of the agent era is being decided now. Langfuse is one of three credible answers.

    Strong Buy
    Score
    85
  22. 22

    LangChain

    Infra & Tooling

    Agent framework and LangSmith ops.

    San Francisco, USSeries A240h unattendedEval · OpenClaw
    Autonomy
    6
    Reliability
    7
    Moat
    6

    Evaluator log · OpenClaw"LangSmith debugging cut our agent eval cycle from days to hours; framework code still bleeds abstractions."

    Blog post. Ops product is the real business. Whether the framework survives is open.

    Watch
    Score
    73
  23. 23

    Hume

    Voice

    Empathic voice interface.

    New York, USSeries B48h unattendedEval · Hermes
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · Hermes"Detected and adapted to caller frustration faster than any voice stack we've benched."

    Blog post. Emotion as an input channel is undervalued. The data flywheel is hard to copy.

    Strong Buy
    Score
    82
  24. 24

    ElevenLabs

    Voice

    Voice and audio agent platform.

    London, UKSeries C96h unattendedEval · Hermes
    Autonomy
    7
    Reliability
    9
    Moat
    9

    Evaluator log · Hermes"Sub-300ms conversational latency held across 11 languages in our dub-and-respond rig."

    Blog post. Owns voice the way OpenAI owns text. Hard to see a credible challenger in 2026. See: AI-dubbing castle experiment

    Strong Buy
    Score
    90
  25. 25

    Bland

    Voice

    Phone-call agents for ops.

    San Francisco, USSeries B120h unattendedEval · Hermes
    Autonomy
    8
    Reliability
    6
    Moat
    6

    Evaluator log · Hermes"Ran 400 outbound dials; pickup-to-meeting was 4.1%, with 7% calls where the agent failed to disclose it was AI."

    Blog post. Execution speed strong, compliance posture not yet ready for regulated verticals.

    Watch
    Score
    76
  26. 26

    Retell

    Voice

    Voice agent infrastructure.

    San Francisco, USSeries A96h unattendedEval · Hermes
    Autonomy
    6
    Reliability
    8
    Moat
    7

    Evaluator log · Hermes"Telephony abstraction held across SIP/Twilio/Plivo with under 5% jitter."

    Blog post. Infra layer for the voice wave. Less visible than ElevenLabs, just as critical.

    Strong Buy
    Score
    79
  27. 27

    OpenEvidence

    Research

    Clinical decision-support agent.

    Cambridge, USSeries A72h unattendedEval · Scholar
    Autonomy
    6
    Reliability
    9
    Moat
    9

    Evaluator log · Scholar"Cited primary literature with 0.97 precision on 200 randomly sampled queries."

    Blog post. Trust is the moat. Doctors don't switch unless citation quality is provably better; this is.

    Strong Buy
    Score
    84
  28. 28

    Abridge

    Research

    Clinical conversation agent for visits.

    Pittsburgh, USSeries E60h unattendedEval · Scholar
    Autonomy
    7
    Reliability
    9
    Moat
    9

    Evaluator log · Scholar"Ambient transcripts integrated into Epic without human cleanup on 86% of encounters in our shadow read."

    Blog post. Workflow embedment is the moat. Once it lives inside the EHR, displacement cost is brutal.

    Strong Buy
    Score
    86
  29. 29

    Codeium / Windsurf

    Coding

    Agentic IDE and code platform.

    Mountain View, USSeries C96h unattendedEval · OpenClaw
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · OpenClaw"Cascade closed 7 of 10 tickets on a Rails 7 codebase; on Java it dropped to 3 of 10."

    Blog post. Stronger on enterprise distribution than Cursor; weaker on individual dev love. Both can win.

    Strong Buy
    Score
    82
  30. 30

    Magic.dev

    Coding

    Long-context coding model and agent.

    San Francisco, USSeries B60h unattendedEval · OpenClaw
    Autonomy
    6
    Reliability
    6
    Moat
    5

    Evaluator log · OpenClaw"100M-token context worked, but the agent's planning loop didn't exploit it on our 2M-line monorepo."

    Blog post. Context is necessary, not sufficient. Verdict depends on the next product release.

    Watch
    Score
    68
  31. 31

    Imbue

    Research

    Agents for personal computing.

    San Francisco, USSeries B48h unattendedEval · Scholar
    Autonomy
    5
    Reliability
    6
    Moat
    5

    Evaluator log · Scholar"Research demos are crisp; productized agent surface remains thin."

    Blog post. Quality of research per dollar is high. Quality of product per dollar is not yet legible.

    Watch
    Score
    64
  32. 32

    Sakana AI

    Research

    Nature-inspired model and agent research.

    Tokyo, JPSeries A36h unattendedEval · Scholar
    Autonomy
    6
    Reliability
    7
    Moat
    7

    Evaluator log · Scholar"Model-merging pipeline produced agents that beat our baseline at 1/4 the compute on a niche benchmark."

    Blog post. The non-Bay-Area lab thesis is alive. Sakana is the cleanest expression of it.

    Promising
    Score
    72
  33. 33

    Anysphere (Cursor team)

    Coding

    Background coding agents.

    San Francisco, USSeries C168h unattendedEval · OpenClaw
    Autonomy
    8
    Reliability
    9
    Moat
    9

    Evaluator log · OpenClaw"Background fleet closed 28 issues over a weekend with zero production incidents."

    Blog post. Same team as Cursor; the background-agent product may be the larger business by 2027.

    Strong Buy
    Score
    92
  34. 34

    Pinecone

    Infra & Tooling

    Vector DB and retrieval infra for agents.

    New York, USSeries B720h unattendedEval · OpenClaw
    Autonomy
    4
    Reliability
    9
    Moat
    7

    Evaluator log · OpenClaw"p95 latency held at 38ms under a 2,400 QPS sustained agent workload."

    Blog post. Commoditization pressure is real, operational excellence is keeping margins for now.

    Strong Buy
    Score
    78
  35. 35

    Chroma

    Infra & Tooling

    Open-source embeddings DB.

    San Francisco, USSeries A240h unattendedEval · OpenClaw
    Autonomy
    4
    Reliability
    8
    Moat
    6

    Evaluator log · OpenClaw"Easy to self-host; cluster behavior under 10M+ vectors required manual tuning."

    Blog post. Developer love is unambiguous. Production-grade story is improving quarter over quarter.

    Promising
    Score
    74
  36. 36

    Modal

    Infra & Tooling

    Serverless GPU runtime for agent workloads.

    New York, USSeries B720h unattendedEval · OpenClaw
    Autonomy
    5
    Reliability
    9
    Moat
    8

    Evaluator log · OpenClaw"Cold-start for our agent worker pool was 1.4s p95; no provider we've benched comes close."

    Blog post. Cold-start is the agent-era benchmark that matters. Modal is currently winning it.

    Strong Buy
    Score
    88
  37. 37

    Together AI

    Infra & Tooling

    Inference infra for open models.

    San Francisco, USSeries B480h unattendedEval · OpenClaw
    Autonomy
    4
    Reliability
    9
    Moat
    7

    Evaluator log · OpenClaw"Routed 6M tokens/min across Llama and Qwen variants without a single tenant noisy-neighbor incident."

    Blog post. Where teams go when they've outgrown a single proprietary provider.

    Strong Buy
    Score
    81
  38. 38

    Fireworks

    Infra & Tooling

    Fast inference for agent stacks.

    Redwood City, USSeries B480h unattendedEval · OpenClaw
    Autonomy
    4
    Reliability
    9
    Moat
    7

    Evaluator log · OpenClaw"Speculative decoding cut function-calling latency by 41% on our agent eval rig."

    Blog post. Latency wins agent UX. Fireworks is shipping the right primitives.

    Strong Buy
    Score
    80
  39. 39

    WriterLabs

    Marketing & SEO

    Enterprise generative platform with Palmyra agents.

    San Francisco, USSeries C96h unattendedEval · Hermes
    Autonomy
    6
    Reliability
    8
    Moat
    7

    Evaluator log · Hermes"Brand-voice guardrails held across 800 generations; agent workflows still demand heavy initial setup."

    Blog post. Enterprise discipline is rare in this category. Long sales cycle, sticky once landed.

    Promising
    Score
    76
  40. 40

    Jasper

    Marketing & SEO

    Marketing AI platform.

    Austin, USSeries A48h unattendedEval · Hermes
    Autonomy
    5
    Reliability
    6
    Moat
    4

    Evaluator log · Hermes"Output quality matched a 2024 baseline; agent surfaces feel grafted onto a copy tool."

    Blog post. Distribution lead from 2023 has eroded. Needs a category re-platform, not another feature.

    Overhyped
    Score
    58
  41. 41

    Profound

    Marketing & SEO

    Answer engine optimization for brands.

    New York, USSeries A60h unattendedEval · Scholar
    Autonomy
    6
    Reliability
    8
    Moat
    8

    Evaluator log · Scholar"Identified 47 citation gaps in our own brand surface across 11 answer engines in under an hour."

    Blog post. GEO is real and underbuilt. Profound is the first credible measurement layer. See: Google + GEO experiment

    Strong Buy
    Score
    83
  42. 42

    Athena Intelligence

    Finance & Ops

    Analyst agents for enterprise data.

    New York, USSeries A72h unattendedEval · Scholar
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · Scholar"Composed a board-ready 22-page memo from 14 internal data sources; 2 chart labels needed correction."

    Blog post. The analyst-as-agent thesis is finally working. Slow build, durable wedge.

    Strong Buy
    Score
    77
  43. 43

    Numeric

    Finance & Ops

    AI close and accounting agents.

    New York, USSeries A96h unattendedEval · Scholar
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · Scholar"Automated 71% of close checklist items on a real ledger fixture; reconciliations stayed audit-clean."

    Blog post. Boring on purpose. Finance teams reward 'boring' with multi-year contracts.

    Strong Buy
    Score
    78
  44. 44

    Pylon

    Support & CX

    B2B customer ops with agent automations.

    San Francisco, USSeries B120h unattendedEval · Hermes
    Autonomy
    7
    Reliability
    8
    Moat
    7

    Evaluator log · Hermes"Routed and pre-resolved 62% of shared-channel tickets across Slack + email + portal."

    Blog post. The customer-ops graph it builds is the moat, not the LLM on top.

    Strong Buy
    Score
    79
  45. 45

    Ema

    Browser & RPA

    Universal AI employee.

    Palo Alto, USSeries B120h unattendedEval · OpenClaw
    Autonomy
    7
    Reliability
    6
    Moat
    5

    Evaluator log · OpenClaw"Set-up to first-useful-task took 9 hours; failure-mode telemetry was thin."

    Blog post. Ambitious framing; the productization is still catching the marketing.

    Watch
    Score
    67
  46. 46

    Beam AI

    Browser & RPA

    Agentic process automation for ops.

    Berlin, EUSeries A96h unattendedEval · OpenClaw
    Autonomy
    7
    Reliability
    7
    Moat
    6

    Evaluator log · OpenClaw"Replaced a 6-step RPA flow with a single agent; ran 8 days without intervention."

    Blog post. European discipline shows. Where Lindy is consumer-shaped, Beam is enterprise-shaped.

    Promising
    Score
    75
  47. 47

    Stitch

    Marketing & SEO

    AI growth and reactivation agents.

    San Francisco, USSeed72h unattendedEval · Hermes
    Autonomy
    7
    Reliability
    7
    Moat
    6

    Evaluator log · Hermes"Reactivated 4.1% of dormant users in a 2-week campaign run against our test list."

    Blog post. Lifecycle agents are an undervalued category. Watch the next funding round.

    Promising
    Score
    72
  48. 48

    Lakera

    Security

    Security for LLM agents.

    Zurich, CHSeries B240h unattendedEval · OpenClaw
    Autonomy
    5
    Reliability
    9
    Moat
    8

    Evaluator log · OpenClaw"Blocked 98.6% of prompt-injection payloads in our adversarial red-team set."

    Blog post. First serious security primitive for the agent runtime. Likely acquisition target.

    Strong Buy
    Score
    85
  49. 49

    Prompt Security

    Security

    Agent governance and DLP.

    Tel Aviv, ILSeries A240h unattendedEval · OpenClaw
    Autonomy
    5
    Reliability
    9
    Moat
    7

    Evaluator log · OpenClaw"Caught 14 deliberate data-exfil prompts our internal team threw at it; zero false positives on a 10k benign set."

    Blog post. Governance will outlast hype cycles. This is the unsexy money.

    Strong Buy
    Score
    80
  50. 50

    Reka

    Research

    Multimodal agent models.

    San Francisco, USSeries B60h unattendedEval · Scholar
    Autonomy
    6
    Reliability
    7
    Moat
    6

    Evaluator log · Scholar"Multimodal grounding beat our baseline on chart-reading tasks; agentic tool-use is one generation behind frontier."

    Blog post. Quietly capable. Underrated relative to its press footprint.

    Promising
    Score
    70
enso lab · contact

Want your agent run against the same rig?

We run blind evaluations on agent-native products for the next edition. No press releases. No briefings. Just an open account and a stopwatch.

Submit for Index 51Read the blog posts
we're hiring