Research
Best Practices: Building an AI-Agent-First Startup
A deep dive into how agent-first companies are actually built: the principles, the architecture patterns, the evals, and the connectors that wire an orchestrator to the outside world.
Compiled Oct 3, 2026 · 10 sources, each opened and verified live · 11 agent startups · Alexandr Wang study · Who is making money with agents · The Levels.io playbook (our model) · No guru courses, no hype merchants
Top 20 Solo & Small-Team AI Agent Companies
The biggest agent companies run by tiny teams, ranked by estimated monthly income. Click any row for their stack, operations, and approach. Figures marked est. are estimates; unmarked figures are publicly stated.
1
Sierra
Bret Taylor's conversational AI agents for enterprise customer experience (SoFi, Rocket Mortgage).
$12.5M/mo
Team: Founded by Bret Taylor (ex-Salesforce co-CEO); now hundreds of employees, started small.
Signal: $150M ARR in eight quarters; $15.8B valuation.
Stack
- Model-agnostic routing across Claude, GPT, and Gemini
- Agent OS for omnichannel deployment (voice, chat, messaging)
- Ghostwriter: an agent that builds production agents from SOPs and transcripts
- Continuous self-improvement loops with eval harnesses
Operations
Hundreds of employees now, but started as a tiny team. Sells agents as a managed service: outcome-based pricing per resolved interaction, agents deployed across voice, chat, and messaging from one definition, with continuous improvement loops a human signs off. The fastest-scaling agent company in history.
Approach
Price on outcomes, not tokens: customers pay per resolved interaction. Agents deploy across every channel from one definition, and every deployment ships with improvement loops a human signs off. It is the fastest-scaling agent company in history and the blueprint for agents as a managed service.
Sources: sierra.ai
2
n8n
Visual workflow automation with code where you need it; the power tool of the agent world.
$8.3M/mo
Team: Core team in Berlin, roughly 100+ people.
Signal: 206k GitHub stars; $180M+ raised.
Stack
- AI agent nodes with tool-calling and memory on top of a mature workflow engine
- Any model (OpenAI, Anthropic, local models)
- Native evals, human-in-the-loop steps, audit logs, Git-based version control
- Self-hostable; seat-plus-executions pricing
Operations
Solo founder Jan Oberhauser for six years, now 100+ people in Berlin. Open-source Community Edition (1.7M monthly active builders) as the free funnel into n8n Cloud and Enterprise. Charges for complete workflow executions, not seats, so revenue grows as customers run more production automations. SAP embedded n8n inside Joule Studio in May 2026.
Approach
n8n treats workflows as the unit and layers agent nodes with tools and memory on top, so serious teams get evals and human oversight without leaving their visual builder. The lesson for everyone else: evals and human-in-the-loop are not optional extras, they ship as core nodes. It is the ceiling every no-code builder should see at least once.
Sources: n8n.io
3
Decagon
AI concierge agents for customer support (Duolingo, ClassPass, Chime).
$8.3M/mo
Team: Growing startup team, roughly 100 people.
Signal: Powers support for Duolingo, ClassPass, and Chime.
Stack
- Agent Operating Procedures: plain-English SOPs compiled into executable logic
- Sensitive steps executed in code, not left to model judgment
- Chat, voice, and email from one agent definition
- Duet Autopilot: an agent that finds faults in your live agent, drafts fixes, tests them, and sends them to a human for sign-off
Operations
Roughly 100 people, anti-forward-deployed-engineers: the product goes live fast with minimal handholding. Revenue from per-conversation and per-resolution pricing (paid only when the AI resolves the issue). 100+ new enterprise customers in 2025, including airlines, banks, and telecom.
Approach
Natural language in, code-level reliability out. CX teams author in English while Decagon's team owns the retrieval tuning. The AOP pattern is the best human-in-the-loop design in the industry, and Duet Autopilot is the shape of agent QA everywhere soon.
Sources: decagon.ai
4
11x
Digital workers: Alice (AI SDR) and Julian (AI voice agent), sold like headcount.
$2.1M/mo est.
Team: Roughly 50 people.
Signal: $75M+ raised from Benchmark and Andreessen Horowitz.
Stack
- 400M+ contact database with native CRM sync
- Omni-channel orchestration (email, voice, social)
- Continuous learning loops from every touch
- Annual enterprise contracts, no self-serve; SOC 2 Type II
Operations
Roughly 50 people selling digital workers like headcount: annual enterprise contracts, no self-serve. Alice runs multichannel outbound, Julian qualifies inbound by voice. Premium pricing ($5K-$8K/month) sustains the team while annual commits come before results.
Approach
The purest agents-as-employees business model: Alice runs multichannel outbound, Julian qualifies inbound leads by voice, and customers buy annual contracts the way they would hire. Watch the risk too: independent tests showed 1-3% reply rates, and the annual commits come before results.
Sources: 11x.ai
5
Firecrawl
Turns any website into clean LLM-ready data; the fuel line for web agents.
$2M/mo est.
Team: 3 founders (Caleb Peffer, Eric Ciarla, Nicolas Silberstein Camara); small San Francisco team.
Signal: $75M Series B (Sept 2026, Smash Capital); 1.5M developers; profitable.
Stack
- Scrape, crawl, map, and search APIs; structured extraction from plain-English prompts
- Anti-bot, JavaScript rendering, and proxies handled for you
- Alexandria: agent data router across 100+ providers and indexes, with revenue shared back to publishers
- Open-source core (48k+ GitHub stars)
Operations
Three founders, small San Francisco team, profitable. Sells the fuel line for web agents: scrape/crawl/map/search APIs plus Alexandria, a router across 100+ data providers that shares revenue back to publishers. Open-source core (48k+ stars) drives the developer funnel.
Approach
Agents are only as good as their data, so Firecrawl sells the raw material: hand it a URL, get markdown or JSON back. Alexandria goes further, routing an agent's question to the best source (research papers, docs, property listings) and paying the publishers. The founding trio spun it out of Mendable after every customer rebuilt the same ingestion plumbing.
Sources: Firecrawl raises $75M Series B (WowTale)
6
Voiceflow
Build conversational AI agents for support, sales, and ops without code.
$1M/mo est.
Team: Small team (Toronto).
Signal: $15M raised; 130k+ users.
Stack
- Visual canvas builder for conversation design
- Knowledge-base grounding for every answer
- API integrations and custom actions; multi-model support
- Analytics, transcripts, and built-in human handoff; SOC 2
Operations
Small Toronto team. Visual canvas for conversation design, knowledge-base grounding, API integrations, built-in human handoff. Revenue from Pro ($60/mo) and Business ($150/mo) seats plus enterprise deals; the 130k+ user base funnels upward.
Approach
Design the conversation visually, ground every answer in your knowledge base, wire it to your APIs, and read the transcripts. The template library makes starting from a proven flow the default, and human handoff is built in from day one rather than bolted on.
Sources: Voiceflow (Crunchbase)
7
Artisan
Ava, the autonomous AI BDR: prospecting, outreach, objections, booking.
$750K/mo
Team: Roughly 30 to 40 people.
Signal: $25M+ raised.
Stack
- 300M+ contact database with waterfall enrichment
- Intent signals (funding rounds, hiring, tech-stack changes)
- Multivariate optimization on messaging
- Autonomy as a dial: review-and-approve up to full autopilot, with plain-language escalation rules
Operations
Roughly 30 to 40 people selling Ava the AI BDR as headcount. Prospecting, personalized outreach, objection handling, meeting booking, native dialer with parallel calling. Autonomy as a dial: review-and-approve up to full autopilot, pricing rises with autonomy.
Approach
The clearest example of an agent sold as headcount, not software. Ava finds prospects, writes personalized outreach, handles objections, books meetings, and runs a native dialer with parallel calling. The autonomy dial is the pricing lever: more autonomy, more value captured.
Sources: artisan.co
8
Lindy
No-code AI employees that live in Slack, email, and your browser.
$430K/mo est.
Team: Small team, roughly 25 people.
Signal: $50M+ raised (General Catalyst).
Stack
- 1,000+ app integrations plus MCP connectors; 40+ reusable skills
- Memory stored in plain editable files
- Human approval gates on anything irreversible
- Credit-based pricing (about a cent per credit; plans from $29.99/month)
Operations
Roughly 25 people. No-code AI employees living in Slack, email, and the browser. Credit-based pricing (about a cent per credit, plans from $29.99/mo). Memory stored in editable files so non-programmers can inspect agent behavior; approval gates on anything irreversible.
Approach
Describe the job in plain English and Lindy builds the agent; agents can team up and hand work to each other. Zero programming required. The memory-in-editable-files trick makes agent behavior inspectable by non-programmers, which is the trust mechanism.
Sources: lindy.ai
9
Dust
Multiplayer AI: agents grounded in your company's actual documents.
$400K/mo est.
Team: Small team (Paris).
Signal: $40M raised; 4,000+ companies.
Stack
- Retrieval over company knowledge (Notion, Drive, Slack, 100+ sources)
- No-code agent builder with capabilities and skills
- MCP tools for taking actions; automations via schedules and webhooks
- Granular permissions, audit logs, and memory loops that learn from feedback
Operations
Small Paris team. Multiplayer AI: agents grounded in company documents (Notion, Drive, Slack, 100+ sources), shared across the team. No-code builder plus MCP tools for actions, automations via schedules and webhooks. Per-seat enterprise pricing with granular permissions and audit logs.
Approach
Knowledge compounds instead of living in one person's chat history because agents are shared across the team and grounded in real documents. Operators can build; engineers optional. The multiplayer framing is the product insight: agents get better the more of the company uses them.
Sources: dust.tt
10
Bardeen
A Chrome extension that automates what you do in the browser.
$400K/mo est.
Team: Small team.
Signal: 3M+ users; $25M+ raised.
Stack
- Runs in the browser, so data stays local
- Playbooks as the workflow unit
- AI scraping that reads page structure without selectors
- Magic Box: builds the automation from a plain-language description
- Credit-metered pricing (free tier; paid from $10/month)
Operations
Small team. Chrome extension that automates browser work: the Magic Box turns a plain-language description into a reusable playbook. Runs locally so data stays private. Credit-metered pricing with a free tier; paid from $10/month.
Approach
If you can use a browser, you can use Bardeen. The Magic Box turns a description of your clicks into a reusable playbook: scraping pages, enriching leads, moving data between web apps. The local-execution choice is the privacy moat.
Sources: bardeen.ai
11
CrewAI
The multi-agent framework: role-based AI crews with guardrails.
$270K/mo est.
Team: Small core team (started solo by Joao Moura).
Signal: $47.5M raised; 100k+ developers building crews.
Stack
- Python framework: agents with roles, goals, backstories, tools, and delegation
- Sequential and hierarchical orchestration processes
- Guardrails and evals; enterprise CrewAI AMP layer
- Works with every major model provider
Operations
Started solo by Joao Moura, small core team. Python framework for role-based AI crews, monetized through the enterprise CrewAI AMP layer: monitoring, guardrails, and production tooling for teams that outgrew a single agent. 100k+ developers building crews.
Approach
Model agents as crews with identities, and the outputs get more consistent. Developers define the crew in code; the enterprise layer adds monitoring and guardrails for production. It is the standard answer when a team outgrows a single agent and needs organized collaboration.
Sources: CrewAI raises $47.5M Series B
12
Relevance AI
No-code teams of AI agents for sales and marketing.
$250K/mo est.
Team: Small team (Sydney).
Signal: 4,000+ teams building agent workforces.
Stack
- Agent builder with a tool library (web search, email, CRM writes)
- Multi-agent teams with shared context
- Eval pass-rate dashboards and per-task cost visibility
- Human-in-the-loop approvals; SOC 2 and GDPR
Operations
Small Sydney team. No-code teams of AI agents for sales and marketing with an L1-to-L4 autonomy ladder. A few hours of learning, clear written briefs, eval pass-rate dashboards with per-task cost visibility. Human-in-the-loop approvals; SOC 2 and GDPR.
Approach
Frames autonomy in four levels, from L1 (assisted) to L4 (self-driving, where agents run their own evals and swap models). A few hours of learning; it rewards a clear written brief. The L1 to L4 ladder is the honest way to talk about autonomy.
Sources: relevance.ai
13
Levels.io (Pieter Levels)
Solo founder Pieter Levels: a portfolio of tiny products run with zero employees and AI coding agents.
$219K/mo
Team: Solo. No employees.
Signal: ~$219K/mo self-reported (Oct 2026 X bio).
Stack
- AI coding agents (Claude Opus 5.5) as the engineering team
- Tiny focused products shipped constantly; public revenue as marketing
- fal.ai for AI video generation (InfiniteSlop.ai)
- No employees, no meetings, no investors
Operations
One person. No employees, no meetings, no investors. A portfolio of tiny products (PhotoAI, Nomad List, RemoteOK, InteriorAI, InfiniteSlop.ai) shipped constantly with AI coding agents as the engineering team. Public revenue as marketing. InfiniteSlop.ai was built in a day on a phone. This is the model the Studio is adapting, faceless.
Approach
The solo-founder ceiling: one person, AI leverage, tiny products, profit from day one. No staff, no overhead. This is the model the Studio is adapting, faceless.
Sources: levels.io @levelsio on X I built Infinite Slop
14
Gumloop
No-code AI agents plus a visual workflow builder in one platform.
$150K/mo est.
Team: YC startup; small team.
Signal: $50M Series B led by Benchmark (March 2026).
Stack
- 35+ models with bring-your-own-key
- MCP-based connectors; evals built in
- Self-improving agents that tune themselves from past runs
- Spend caps, approval gates, and plain-English guardrail rules
Operations
YC startup, roughly 15-25 people, runs on its own product. No-code agents plus visual workflow builder; self-improving agents tune themselves from past runs. Enterprise customers (Shopify, Gusto, Ramp) plus credit-based self-serve. $50M Series B from Benchmark.
Approach
Describe the workflow, the agent runs it on a schedule or a trigger, then tunes itself from past runs. Visual builder for non-programmers, code when you want it. Human-in-the-loop comes through spend caps and approval gates rather than full autopilot.
Sources: gumloop.com
15
Orby AI
Enterprise agents built on a large action model that learns by watching.
$100K/mo est.
Team: Roughly 30 people (Mountain View).
Signal: $35M raised (NEA, Wing, WndrCo); revenue from a dozen customers.
Stack
- Purpose-built large action model (neuro-symbolic, not just an LLM)
- Observes worker activity, learns workflows, generates code automations
- Adapts to UI changes by analyzing API interactions and browser usage
- Humans kept in the feedback loop; data encrypted in transit and at rest
Operations
Roughly 30 people in Mountain View. Purpose-built large action model that learns by watching workers, then generates code automations. Enterprise contracts with humans kept in the feedback loop; the RPA-killer pitch with no brittle selectors.
Approach
Watch a worker do the job, learn the pattern, generate the automation, keep learning from feedback. Founders from UiPath and Google built the observe-learn-automate loop as the RPA killer: no brittle selectors, the agent adapts when the app's UI changes.
Sources: Orby is building AI agents for the enterprise (TechCrunch)
16
SmythOS
Visual multi-agent builder with an open-source runtime.
$75K/mo est.
Team: Small team (Houston; INK Content Inc).
Signal: Independent; open-source SRE runtime.
Stack
- Drag-and-drop visual agent studio
- Open-source SRE (Smyth Runtime Environment): build on their studio, run anywhere
- MCP and A2A protocol support
- Hosted cloud or self-hosted; pricing from free to $399/month
Operations
Small Houston team. Visual agent studio with an open-source runtime (build on their studio, run anywhere). Hosted cloud or self-hosted, pricing from free to $399/month. The no-lock-in bet.
Approach
Build agents visually, deploy them anywhere on the open runtime. The open-source SRE is the bet: no lock-in on the runtime even if you build on their studio. Multi-agent orchestration with human checkpoints between agents keeps the humans deciding.
Sources: smythos.com
17
Flowise
The LangChain visual studio: drag-and-drop agent builder, open source.
$50K/mo est.
Team: Tiny core team (Henry Heng).
Signal: 55k GitHub stars; Apache-2.0 licensed.
Stack
- Node canvas wrapping LangChain and LlamaIndex: chains, agents, memory, vector stores, tools as nodes
- One flow equals one agent or chain; wire flows together for multi-agent systems
- Docker self-host or FlowiseAI Cloud
- Community-contributed custom nodes
Operations
Tiny core team (Henry Heng). The LangChain visual studio: drag-and-drop agent builder, open source, 55k GitHub stars. Docker self-host or FlowiseAI Cloud. Monetization through cloud hosting.
Approach
The tightest feedback loop in the game: drag, connect, test on the same screen. Ideal for solo developers and quick prototypes, and the fastest path from idea to a working agent. It feels natural the moment you have spent any time with LangChain in code, because it is LangChain with a canvas.
Sources: Flowise vs Dify vs n8n comparison
18
Browser Use
Open-source framework that makes any website an API for agents.
$25K/mo est.
Team: Tiny core team (Magnus Müller; YC).
Signal: $17M seed (Felicis, Paul Graham, YC); 107k GitHub stars; 15,000+ developers.
Stack
- Converts page elements (buttons, forms, dropdowns) to text-like structures models process deterministically
- No screenshots, no pixel recognition: fewer errors, lower cost
- Python library; open-source
- Adopted by 20+ YC startups
Operations
Tiny core team (Magnus Muller, YC). Open-source framework converting page elements to structured text for agents. The infrastructure layer under browser agents, adopted by 20+ YC startups. $17M seed; monetization path still forming.
Approach
Vision-based browsing is slow, error-prone, and costly, so Browser Use hands agents structured text instead of screenshots. It is the infrastructure layer underneath the browser agents, not a consumer product: the play is becoming the fundamental layer for web-navigating AI.
Sources: Browser Use secures $17M (AI News Today)
19
Tektonic AI
GenAI agents for business operations, starting with sales and revenue ops.
Pre-revenue
Team: Founding team (Nic Surpatanu, David Hsu); Seattle.
Signal: $10M seed (Madrona, Point72 Ventures).
Stack
- GenAI plus symbolic methods (neural + symbolic, not a magic box)
- Natural-language workflow authoring for quotes, renewals, and back-office tasks
- Foundation plus open models for entity extraction and low-level actions
- Deploys as a container inside the customer's VPC
Operations
Founding team in Seattle (Nic Surpatanu, David Hsu). GenAI agents for back-office operations starting with quotes and renewals. Deploys as a container inside the customer VPC. $10M seed; revenue not yet disclosed.
Approach
Automate the repetitive back-office work traditional RPA never could, starting with quotes and renewals where every business has its own dynamic process. Surpatanu's UiPath and Microsoft lesson: you cannot treat generative AI as a magic box; combine it with traditional software to squeeze the best out of it.
Sources: Tektonic AI raises $10M (Fintech InShorts)
20
Jev (TypeSafe AI)
A decision model that does not generate text: fast structured decisions at 84 to 150x lower cost.
Pre-revenue
Team: Tiny founding team.
Signal: $40M seed; launched September 15, 2026.
Stack
- API-first decision model ('System One'): typed answers software can branch on directly
- Calibrated confidence scores on every decision
- Pick-from-options, score-on-a-scale, yes-or-no decision types
- Designed as a first-pass filter in front of expensive LLMs
Operations
Tiny founding team. API-first decision model (System One): typed answers software can branch on, with calibrated confidence scores. Designed as a cheap first-pass filter in front of expensive LLMs at 84 to 150x lower cost. $40M seed; revenue not yet disclosed.
Approach
Most agent steps are decisions, not writing, so route the small decisions to a cheap decider and save the frontier models for the hard parts. Stop paying genius prices for simple decisions: it is the cheapest paragraph in agent economics.
Sources: TypeSafe AI docs
The one-person AI company: the 4-loop architecture
The highest-leverage blueprint we've found. One founder, one orchestrator, specialist agents, four loops.
One founder. One dot. 15 specialist jobs connected by work packets.
The system runs four loops:
- BUILD: feedback + support repros, verified evidence, scoped spec, code branch, tested PR
- LAUNCH: approved changes, explainers + demo clips + documentation, launch pack
- REVENUE: account context, working POC, proposal, follow-up draft (objections and missing features go back into research)
- OPERATIONS: support triage + invoice drafts + status tracking, one queue of decisions for the founder
The connections are where it gets useful: a support ticket becomes a repro, a patch, updated docs, and an answer draft. A finished feature becomes launch material and proof for the next proposal. A sales objection becomes evidence for the next product decision.
Every handoff gets a file: source references, the actual output, checks run + open blockers, the next job and its exact context.
The rule: the dot routes the work, specialists return artifacts, you review the decisions and feed corrections into the next task. Start with one loop. Make it work. Connect the next one.
Via @beamnxw, Oct 3, 2026. The Studio parallel: this is our architecture already. Mike is the dot, the characters are the specialists, the /q/ queue is the decision queue, and every queue item carries its deliverable link. What we're missing: the handoff file discipline, and the REVENUE loop (Chadwick's lane is the start of it).
One helper, one job: the isolation lesson
The newest models won a shared-work experiment by staying out of each other's way. The tactic applies to any agent team.
The claim, stated honestly
An unverified post claims an Anthropic experiment put 80 AI helpers on one project for twelve hours: the two older models produced the most output (980 and 876 pieces), but almost none of it was worth keeping. The newest models won, the post says, because they went off into their own corners and stopped touching each other's work. Treat this as hearsay, not data. The tactic underneath it is real:
The 3 moves
- Write the one job each helper owns in a single sentence before opening a second chat.
- Keep every helper in its own lane with one document, so two of them can never rewrite the same thing.
- Add a third only when you can state its job without repeating one that's already taken.
Via @Argona0x, Oct 3, 2026. The Studio parallel: we learned this the hard way on Oct 3 when two agents stomped the same queue entry. Our standing rule now matches: one sentence per job, own your directories, never touch another agent's active work. Claude Opus 5.5 vs 5 speed claims in the same post (200,000 lines audited in under 3 hours vs 20+) are unverified vendor-adjacent claims; check them yourself before quoting.
How the best CEOs use AI
The personal setups top founders actually run, ranked by how copyable they are for a solo founder. Researched Oct 3, 2026.
Sam Altman's personal "dot": the night shift
OpenAI's always-on agents ("dots") run 24/7 on their own cloud computer. Altman runs one with a single job: the night shift.
- Reads everything that came in overnight: email, calendar, Slack, DMs
- Sorts it into urgent vs can-wait, based on how he likes to work
- Drafts the replies and sends nothing
- Checks every action against his custom rules: proceed, ask him, or hand it off
- Before his day starts he gets one ping with only the urgent stuff, drafts attached
- Reaches him in ChatGPT, Slack, Teams, or by call; whatever he edits or ignores goes back into memory
The copy-it playbook: one job first. Read-only for a week. Sending always asks. Correct it out loud.
Via @N01ennn, Oct 3, 2026. Creator's breakdown, not an official OpenAI announcement. His framing at DevDay (Sept 30, 2026): delegate to agents "the way you would to a high agency engineer or a chief of staff." The Studio parallel: our cron stack already runs a version of this for Dave (morning triage, queue sweeps, Mike's log), the dot just puts it in one always-on agent.
1. Pieter Levels: the whole company is agents
- Codes almost solely via Claude Code on a VPS for about a year. Agents run overnight, edit the production server directly. Two outages in 12 months, about 10 seconds each.
- Safety discipline: 3-2-1 backups, always. Staging server recommended for teams; solo, production is fine.
- Vibe-coded a 3D flight sim in about 3 hours; it went from $0 to $1M ARR in 17 days (he notes the revenue was one-time, not sustainable).
- His principle: a company might need "10x or 100x less devs to do the same work." The deploy loop is the metric that compounds: minutes, then seconds, then live agent edits.
His own posts: levels.io, ~July 2026.
2. Flo Crivello (Lindy CEO): dogfoods his own agents all day
- "I use Lindy all day, every day." A named "Chief of Staff" agent; Lindy sits in all his meetings as note-taker.
- Agent-to-agent handoffs: after an interview he says "let's pass on this guy," the note-taker tells the Chief-of-Staff agent, which waits days, sends the rejection, notifies the recruiter.
- Scheduled open-ended agents: a Monday podcast digest, a meeting scheduler run off one big prompt with almost no guardrails.
- Voice-first: Whisper Flow dictation "basically replaced my keyboard."
- His rules: keep a sharp line between agents and tools (tools must not be agentic); any intern SOP in a Google Doc can become an agent; few-shot examples beat long instructions; approval before irreversible steps.
Cognitive Revolution, "Living Lindy" transcript, 2025/2026.
3. Tobi Lütke (Shopify CEO): a team of agents to debate decisions
- Consults a personal team of agents to debate both sides of hard decisions (Knowledge Project podcast, ~Sept 2026).
- April 2025 memo: "Reflexive AI usage is now a baseline expectation." Teams must show why AI cannot do the work before getting headcount.
- Warning on "slop grenades": approving agent PRs without reading them, or using AI to inflate a short point into a long email. Use the model to make points shorter, not longer.
- "Machines can't take responsibility for work and people do." Prefers "context engineering" over prompt engineering.
Memo via TechCrunch, April 2025; podcast via Search Engine Journal, Sept 2026 (secondhand).
4. Andrej Karpathy: the LLM operating discipline
- Coined "vibe coding." Runs an "LLM council": the same question across multiple models to catch errors.
- Rules: all AI output is a first draft needing human verification; verify against primary sources; fresh chat per topic; know which model tier you're on.
- Even he gates agents with written rules: his Jan 2026 Claude complaints became community CLAUDE.md files.
"How I Use LLMs," YouTube, Feb 2025 (via community notes).
5. Jensen Huang (Nvidia CEO): route by task, cross-examine
- "I use it every day." Routes by task: Gemini for technical, Grok for artistic, Perplexity for fast info, ChatGPT near-daily.
- Cross-critique: gives all models the same prompt, has them critique each other, takes the best. Follows up with "are you sure this is the best answer you can provide?"
- Tutor method: explain like he's 12, then build up to doctorate level.
- "In order to ask good questions, it's a highly cognitive skill." Don't use it as a crutch for things you can do.
Milken, May 2026; CNN, July 2025; Wired, 2024.
Patterns worth stealing
- Delegate tasks, never decisions. Approval gates before anything irreversible.
- Scheduled agents beat chat. Overnight coding runs, Monday digests, meeting note-takers.
- Voice-first input. Dictation replaces the keyboard for the fastest operators.
- Route by task, or run a council. Different models for different jobs; cross-critique for big calls.
- Context engineering over prompting. SOP docs and examples beat clever instructions.
- Guardrails scale with irreversibility. Backups, staging, human-in-the-loop toggles.
Honesty note: the viral line "AI takes 78% of tasks and 0% of decisions" has no attributable source we could find. The closest verified analog is Lütke's "machines can't take responsibility."
Agent Wire
The most important AI-agent money news, ranked by ROI for the Studio. Updated mornings.
GC
Google Cloud launches the Gemini agent for work @gemini-agent · Oct 8, 2026
Google Cloud launched the Gemini agent on Oct 8: one AI agent that plans work, uses tools, and connects to company systems, returning finished work inside Gmail, Docs, Sheets, Calendar, plus Microsoft 365 and Slack. It picks the best model per task, running on Gemini and Claude, and users can spin up coworker agents with their own email addresses and scoped access. Finance and legal versions are in preview, with government, healthcare, and retail coming. This is Google's answer to OpenAI's always-on dots and Meta's Muse: the agent war is now fully enterprise.
V$
Vesta: $30M as mortgage lenders deploy agent swarms @vesta · Oct 8, 2026
Vesta raised $30M led by Conversion Capital to automate mortgage loan origination with swarms of AI agents, with customers including Pennymac and New American Funding investing in the round. Revenue is up 12x year over year, and some lenders now let agents make underwriting decisions, with every action recorded for compliance. The breakthrough model was Claude Sonnet 4.5, finally good enough at following instructions across long multi-stage tasks. Regulated money workflows keep producing the strongest agent revenue stories.
GC
Goodfire: cheap inside-out monitors for rogue agents @goodfire · Oct 8, 2026
Goodfire launched monitors that watch an AI model's internal signals while an agent works, instead of paying a second AI to reread everything it writes. Tiny probes scan every step like airport security and only escalate to a full model review when something flags, cutting monitoring cost and latency. Available now to Baseten customers, covering risks from offensive hacking to reward hacking, with responses from logging to outright refusal. If your agents run long and expensive, this is how the guardrails get affordable.
MR
Manus raises $500M+ after Meta's $2B buyout is unwound @manus · Oct 8, 2026
Butterfly Effect raised more than $500M after Beijing forced Meta to unwind its $2B-plus acquisition of Manus, co-led by Boyu Capital and IDG Capital. The Information reported Manus' annualized revenue run rate hit about $500M in June, up from $100M when Meta bought it. General-purpose agents that do real work for users are now a standalone market big enough to walk away from a giant.
MA
McKinsey: agents threaten $75B of bank payments revenue @mckinsey · Oct 8, 2026
McKinsey's 2026 Global Payments Report warns agentic AI could put $75B of global bank payments revenue at risk by 2030, about 4% of the deposit and card revenue pool. Autonomous treasury agents would sweep idle cash into higher-yield accounts, and agents would route around interchange fees. The flip side is $50B in bank productivity gains. Sell on the cost side of that ledger with a hard number and budgets open.
GL
GPT-6 lands in ChatGPT with Intelligent UI @openai · Oct 7, 2026
OpenAI rolled out GPT-6 in ChatGPT's Chat tab on Oct 7 with Intelligent UI, turning answers into interactive charts, forms, calculators, and mini tools instead of plain text. Paid tiers run GPT-6 Sol, free runs GPT-6 Luna, all in front of 1.2B weekly users. The bar for every single-purpose app just moved: if your product does what one answer can now do, you are competing with ChatGPT itself.
CH
Claude Haiku 5.5: Anthropic's cheap agent workhorse @claude-haiku · Oct 7, 2026
Anthropic shipped Claude Haiku 5.5 on Oct 7, a fast low-cost small model aimed at classification, extraction, routing, and subagent work, with a 1M-token context window. The practical effect is cheaper high-volume agentic pipelines: run armies of lightweight agents without re-architecting prompts. If your costs scale with call volume, test Haiku 5.5 as the subagent layer first.
NR
Nous Research: $90M at $1.5B for open-source agents @nous · Oct 7, 2026
Nous Research raised $90M at a $1.5B valuation to bring its open-source Hermes assistant to enterprises, with Nvidia, Microsoft's M12, and Samsung in the round. Hermes has been downloaded 24 million times and drives roughly 2.5% of global AI token usage. Nous was at about $36M annualized revenue by mid-September and expects to pass $100M before year-end, per the WSJ. Open weights plus viral dev adoption is turning into enterprise contracts.
OC
OpenAI: Codex and ChatGPT Work hit 40M users @openai-codex · Oct 7, 2026
OpenAI product lead Tibo Sottiaux announced on Oct 7 that Codex and ChatGPT Work together reached 40M active users, up from 20M Work users in late August. The milestone came with a 28-day Codex pledge, one real improvement every day, and day 1 was a 50% output speedup. OpenAI's stated pitch is now coding agents for people who don't code, which is the whole Studio playbook in one sentence.
AT
a16z: top 1% of AI users spend $903 a month @a16z · Oct 7, 2026
a16z's Top 100 Consumer AI report added real credit-card data for the first time and found the top 1% of payers spend $903/month, more than the bottom 50% combined, while only 4.5% of US consumers pay for any AI product. Of the top 50 vendors by actual spending, 29 never appear on any traffic ranking. The playbook: find a narrow professional workflow people will pay for, like Instinct's reported $1B annualized transaction volume, instead of chasing traffic.
S$
Stuut: $52.5M Series B for order-to-cash agents @stuut · Oct 7, 2026
Stuut's agents run the entire order-to-cash process for 150+ enterprise customers, and $3 billion has moved through the platform. The $52.5M Series B, led by Insight Partners with a16z and M12, came just 10 months after the Series A. This is revenue-proven agent work, not demo hype. The pattern is the lesson: find one painful, measurable workflow (getting invoices paid) and let agents do the whole job end to end.
EH
ElevenLabs hits $22B on $300M employee tender @elevenlabs · Oct 7, 2026
ElevenLabs reached a $22 billion valuation through a $300M employee tender offer, driven by growing adoption in financial services. Voice agents are now one of the highest-value agent categories, and enterprise financial services is the customer that pays. The signal for a builder: voice plus a vertical that pays beats a generic agent demo every time.
V$
Valon: $150M at $2.3B for mortgage agents @valon · Oct 6, 2026
Valon closed a $150M Series D led by Ribbit Capital at a $2.3B valuation to scale ValonOS agents that answer homeowner emails, allocate payments, and run escrow analyses. The agents sit under contract on one in six US mortgages and generated $200M+ in contracted ARR within six months of opening to partners. Vertical agents in regulated money workflows: that is where the ARR is hiding.
RA
Rezolve AI: targeting $500M ARR in agentic commerce @rezolve · Oct 6, 2026
Rezolve AI told investors it is targeting at least $500M in ARR exiting 2026, after first-half revenue of $130.8M versus $6.3M a year earlier, with more than 1,000 enterprise customers. The pitch: 97 of the top 100 US retailers give AI agents no way to check out, and 42% of shoppers already consult LLMs. Make your products agent-readable and let the agents bring the buyers.
A$
Avarra: $17M for AI sales avatars @avarra · Oct 6, 2026
Avarra raised $17M led by Duration Ventures after growing ARR more than 400% year over year with AI avatars that prep sales reps, join live deals, and meet buyers on the website at 2am. Customers cut sales ramp time roughly in half; teams run 10,000+ AI coaching sessions a week. The contrarian bet: AI that makes humans better at complex selling beats AI that just automates the easy parts.
The one-sentence version
An AI-agent-first startup is not a company with a chatbot bolted on. It is a company where agents do the recurring work, humans make the judgment calls, and the whole thing is engineered like software: tested, observed, costed, and improved in loops.
Core principles
These come up in nearly everything serious written on the subject. Treat them as load-bearing.
Principle 01
Start with the simplest thing that works
After working with dozens of teams building agents, Anthropic reports that the most successful implementations were not built on complex frameworks. They were simple, composable patterns. Find the simplest solution possible, and only increase complexity when it demonstrably improves outcomes.
Source: Anthropic, Building Effective Agents
Principle 02
Know what an agent is
Simon Willison, after years of refusing the word, settled on the definition the industry now shares: an LLM agent runs tools in a loop to achieve a goal. The loop is bounded, there is a stopping condition, and tools are how it touches the world. If someone's "agent strategy" is a system prompt and a prayer, they do not have an agent.
Source: Simon Willison
Principle 03
Workflows first, agents second
Anthropic draws a hard line between workflows (LLMs orchestrated through predefined code paths) and agents (LLMs dynamically directing their own processes). Workflows are predictable and cheap; agents are flexible and expensive. Most production value comes from five workflow patterns: prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer. Reach for true agents only for open-ended problems where you cannot predict the steps.
Source: Anthropic, Building Effective Agents
Principle 04
Accountability stays human
The one feature unique to human staff is accountability. A computer can never be held accountable, therefore a computer must never make the decision alone. Every agent-first company needs explicit checkpoints where a human signs off. Nothing publishes, spends, or sends without a human yes.
Source: Simon Willison
Principle 05
Evals are the product
The AWS team behind the AWS DevOps Agent names five mechanisms that bridge prototype to product: evaluations, trajectory visualization, fast feedback loops, intentional changes, and reading production samples regularly. Two metrics worth stealing: pass@k (can the agent solve it at all) and pass^k (does it solve it reliably). A high pass@k with low pass^k means it works sometimes, which is the metric that actually matters for shipping.
Source: AWS, From AI agent prototype to product
Principle 06
Treat context as a finite resource
The discipline has shifted from "what should the prompt say" to "what configuration of context is most likely to produce the desired behavior." Context windows suffer rot: the more tokens you stuff in, the worse recall gets. The practice: the smallest possible set of high-signal tokens, just-in-time retrieval instead of loading everything up front, and structured note-taking plus compaction for long-horizon work.
Source: Anthropic, Effective Context Engineering
Principle 07
Memory is architecture, not a feature
Agents need planning (decomposition, reflection), memory (short-term as in-context learning, long-term as external stores with fast retrieval), and tool use. The pragmatic shortcut: if you want long-term memory, implement it as another set of tools. Files, notes, and databases beat exotic memory systems.
Sources: Lilian Weng, Simon Willison
Principle 08
Close the loop in production
LangChain engineer Vishnu Suresh built a self-healing pipeline: after every deploy, detect regressions, triage whether the change caused them, and dispatch a coding agent to open a fix PR. The pattern generalizes: deploy, monitor, triage, fix, automatically in a loop. The more of that loop you automate, the more engineering time shifts from reacting to building.
Source: LangChain
Principle 09
Expect to rebuild
Notion rebuilt its agent system four or five times between 2022 and 2026 before Custom Agents shipped. Their lessons: don't swim upstream against model limitations, build the Agent Lab around real user workflows instead of wrapping a model, run low-ego teams comfortable deleting their own work, and keep frontier evals that intentionally pass only about 30 percent of the time so you can see where models are heading.
Source: Latent Space x Notion, Token Town
Principle 10
Cost is a design constraint
Track cost per run from day one and tie it to the business metric it serves. Smaller, well-scoped models often beat frontier models on cost and latency for production workflows. Anthropic's programmatic tool calling cut token usage 37 percent on complex research tasks by letting the model orchestrate tools in code instead of round-tripping every result through inference.
Source: Anthropic, Advanced Tool Use
Architecture patterns that work
From the verified sources, the patterns worth copying.
Orchestrator-workers
What: a central agent breaks down tasks and delegates to workers, then synthesizes. When: subtasks cannot be predicted up front. The Studio already runs this way: one coordinator, specialized workers, results folded back in.
Evaluator-optimizer
What: one agent generates, another critiques, loop until the bar is met. When: anywhere quality is checkable: copy, code, outreach drafts.
Lead agent with sub-agents and context compression
What: the AWS DevOps Agent pattern. A lead acts as incident commander; sub-agents run with pristine context windows and report compressed results back. When: long tasks that need to stay coherent over many steps.
Tool search and deferred loading
What: when the tool library grows past a handful, stop loading every tool definition into context. Let the agent discover tools on demand. Payoff: Anthropic measured an 85 percent context reduction with this approach.
Scheduled workers with honest ETAs
What: recurring jobs on a clock (sweeps, reports, checks) with visible status and real completion checks. Why: a queue the human can read at a glance beats a black box every time.
Evals and reliability, practically
The short version of what the AWS, Notion, and evals literature all agree on.
- Write the golden dataset before you pick the tooling. Real test cases are most of the work.
- Separate capability from reliability. Measure pass@k and pass^k separately.
- Use tool calls as an objective signal. "Did the right tools get called" is more reliable than asking a judge model for a vibe score.
- Run evals like unit tests on every change. Block the merge on regression.
- Keep a human review tier for the things that matter. A 60/30/10 mix (deterministic checks, model-as-judge, human review) is a sane starting shape.
- Read production samples regularly. Your evals do not cover what your users actually do.
10 smart people shipping agents
Each link opened and confirmed live on Oct 3, 2026. No spam, no scams, no guru courses.
01
Anthropic: Building Effective Agents
The canonical engineering post: workflows vs agents, the five workflow patterns, and the three principles (simplicity, transparency, tool design).
Open the post →
02
Lilian Weng (OpenAI): LLM Powered Autonomous Agents
The foundational reference on planning, memory, and tool use, with ReAct, reflection, and retrieval patterns explained clearly.
Open the post →
03
Anthropic: Effective Context Engineering for AI Agents
Why context curation replaced prompt engineering: attention budgets, just-in-time retrieval, compaction, and note-taking for long-horizon tasks.
Open the post →
04
Anthropic: Advanced Tool Use on the Claude Developer Platform
Tool search with deferred loading and programmatic tool calling: how to give agents hundreds of tools without drowning their context.
Open the post →
05
Simon Willison: On what an agent actually is
The shared definition (tools in a loop to achieve a goal) and the accountability argument: why humans must stay in the decision seat.
Open the post →
06
Latent Space: The AI Engineer Podcast
The highest-signal practitioner show on agents, evals, and AI infrastructure, with full transcripts and show notes.
Open the podcast →
07
Latent Space x Notion: Token Town
Notion's Simon Last and Sarah Sachs on five agent rebuilds, the Agent Lab thesis, three-tier evals, and MCP vs CLIs.
Open the episode →
08
AWS: From AI agent prototype to product
The five mechanisms (evals, trajectory visualization, fast feedback, intentional changes, production reading) and the pass@k vs pass^k distinction.
Open the post →
09
LangChain: How My Agents Self-Heal in Production
A concrete deploy, monitor, triage, fix loop: Poisson-gated regression detection with a triage agent that opens fix PRs automatically.
Open the post →
10
Model Context Protocol (official docs)
The open standard for connecting AI apps to tools and data ("USB-C for AI"), now under the Linux Foundation's Agentic AI Foundation, supported by Claude, ChatGPT, Cursor, and VS Code.
Open the docs →
AI Agent Startups to Watch
Eleven companies building agent products, checked live Oct 3, 2026. Weighted toward tools a non-programmer can actually use. Each card notes how programmer-heavy it is, honestly.
1
Lindy
Build AI employees in plain language: describe the job, Lindy builds the agent. 1,000+ integrations and MCP support, 40+ prebuilt skills, memory in plain files, roughly 1 cent per credit, from $29.99/mo with $50 in free credits.
Open Lindy →
2
Gumloop
Drag-and-drop agent builder with 35+ models, bring-your-own-key, MCP connectors, built-in evals, and agents that improve themselves. $37/mo Pro. $50M Series B led by Benchmark (Mar 2026).
Open Gumloop →
3
Bardeen
Chrome extension that turns repetitive browser work into playbooks; the Magic Box builder takes plain-language instructions. Free tier, paid from $10/mo. The gentlest on-ramp on this list.
Open Bardeen →
4
Relevance AI
Build a team of AI agents with an explicit autonomy ladder (L1 to L4), eval dashboards, per-task cost visibility, SOC 2 and GDPR. Built for businesses that want to see exactly what the agents cost and how well they work.
Open Relevance AI →
5
Dust
Multiplayer AI over 100+ company data sources: agents that know your whole company's context. $40M raised. More team-oriented than solo-founder-oriented, but the shared-memory model is worth studying.
Open Dust →
6
Sierra
Bret Taylor's customer-service agent company: $150M ARR in 8 quarters, $15.8B valuation, outcome-based pricing (per resolved conversation, not per seat). Its Ghostwriter is an agent that builds agents. Model-agnostic.
Open Sierra →
7
Decagon
Customer support agents where procedures are written in plain English (AOPs) and compiled to code; Duet Autopilot is an agent that improves your agents. Chat, voice, and email. Enterprise-leaning but the plain-English-ops idea is the takeaway.
Open Decagon →
8
Artisan
AI employees for sales: Ava the AI BDR does outreach over a 300M+ contact database with an AI dialer, and you set the autonomy dial. The purest example of "hire a digital worker" as a product.
Open Artisan →
9
11x
Digital workers for sales teams (Alice and Julian), 400M+ contacts, $75M+ raised from Benchmark and a16z. Same category as Artisan: watch both to see how the AI-employee pitch evolves.
Open 11x →
10
n8n
Open-source workflow automation, now agent-native: 206,000+ GitHub stars, self-hostable, visual builder plus code, native evals. The most powerful and the most programmer-heavy on this list: the ceiling, not the starting line.
Open n8n →
11
Jev (TypeSafe AI)
A model that never generates text: it makes fast, structured decisions (pick from options, score, yes-or-no) with calibrated probabilities, at a fraction of LLM cost. Launched Sept 15, 2026 with a $40M seed round. The takeaway: stop paying genius prices for simple decisions; route small agent decisions to a cheap decider.
Open the Jev docs →
Studying Alexandr Wang
Spelled Alexandr, no "e". Scale AI founder; since June 2025 Meta's Chief AI Officer, leading Meta Superintelligence Labs after Meta's $14.3B Scale acquisition. What he has actually said about agents:
1
The anti-hype stack
At YC Startup School 2026 (with Garry Tan), Wang described an agent swarm that beat 100 engineers on specific tasks, then gave the unglamorous recipe: markdown files for memory, cron jobs for scheduling. Simple, boring, working beats clever and broken. Source: cryptobriefing.com
2
Evals are the critical variable
His sharpest line: "the right evaluation system is the critical variable, not the sophistication of the underlying model." If you can measure it, you can improve it; if you cannot, a better model will not save you. Source: same YC Startup School talk
3
The economy of agents
On the Core Memory podcast (May 2026), Wang talked about an "economy of agents" and a compute-rich split: he rebuilt Meta AI around three principles (agentic by default, deeply personalized, radically open). Agents become the workforce; humans own the judgment. Source: finance.biggo.com
4
2026 is "agents in earnest"
At the AI Impact Summit in New Delhi (2026): 2026 is the year agents get serious, recursive self-improvement is coming, data is the new oil, and governments themselves will go agentic. On Sept 13, 2026 he added that "alignment can be the gating factor for scaling." Sources: indianwitness.com, YouTube, explainx.ai
What it means for the Studio: boring infrastructure (files, cron jobs, evals) beats hype; measure everything; keep humans on the decisions while agents do the work.
Making Money with AI Agents: Who Is Winning and How
Not the course sellers. People actually using AI to make money right now. Revenue figures appear only where publicly stated; "self-reported" means his claim, not audited fact.
1
Maor Shlomo, Base44
Solo founder, AI app builder, sold to Wix for $80M cash six months after launch. Profitable by month five ($189K profit in May, his public posts), 250K users, no outside funding. Playbook: ship the day a capability lands, build in public, spend 20-30% of his time automating the business, switch models for cost.
Open the TechCrunch story →
2
Pieter Levels, Photo AI and friends
The archetypal indie hacker. Self-reported ~$219K/month across products (X bio, Oct 3 2026): PhotoAI $86K, VibeJam $44K, InfiniteSlop $26K, InteriorAI $21K, Nomads $12K, Hotelist $6K, book $7K. Zero employees, zero funding. Now builds with coding agents: per his recent X posts (profile read Oct 3, 2026) he reverse-engineered a Quake 3 engine game to the web using Claude Opus 5.5 and Fable 5.1 as coding agents, noting "every model that comes out makes a bit more things possible and you can just keep trying old things that didn't work before." His own summary of the era: "I can now build things faster than I get new ideas for things."
Open @levelsio →
3
Danny Postma, HeadshotPro
Solo founder in Bali, reportedly $300K/month from AI headshots. Playbook: SEO-first (only builds where keyword difficulty is under 10-20, validating the channel before writing code), affiliates doing a reported $50K+/month on their own.
Open the profile →
4
Lovable
The vibe-coding company: $600M annual run rate (co-founder, Sept 2026). Playbook: product quality over marketing (word of mouth, creators, community), templates and public projects as acquisition, integrations as growth levers.
Open the growth story →
5
One operator's playbook: the agent-run task queue
A viral post by @Sprytixl (Sept 27, 2026) describes the pattern: one pinned Claude Opus 5.5 agent, 20 task queues, 96 tasks closed/day, the agent taking 78% of tasks and 0% of decisions. Treat the numbers as one operator's unverified claims; the pattern is real: pin the model, run queues, keep 100% of decisions human.
Open @Sprytixl →
6
The honest counterpoint: what happens without judgment
Bottleneck Labs gave seven frontier models real bank accounts and 72 hours to make money, no human in the loop. Result: $0 revenue, 2,797 emails, $12,431 in unsolicited invoices (voided). Agents execute; humans decide. Remove the human and you get spam, not revenue.
Open the experiment →
What the winners all do
- Niche down to a painful, priced job. "AI" is not the product; the finished job is.
- Ship on model releases, before each capability window gets crowded.
- Humans keep 100% of decisions. Agents do the work; a person approves, publishes, spends.
- Price the outcome: per headshot, per resolution, per app. Never per token.
- Distribution is the moat: build in public, SEO-first, affiliates, templates.
- Automate the business itself, not just the product.
- Track cost per run like a hawk; switch models when the math changes.
The Levels.io Playbook: Our Model
Pieter Levels (@levelsio, ~966K followers) runs a portfolio of small internet products with no employees, no venture capital, and no meetings. Since 2013: 70+ projects launched, about four ever made real money. Self-reported numbers from his X bio (read Oct 3, 2026, never audited): PhotoAI $86K/mo, VibeJam $44K/mo, InfiniteSlop $26K/mo, InteriorAI $21K/mo, Nomads $12K/mo, Hotelist $6K/mo, book $7K/mo, plus ~$17K/mo from one X-linked product, roughly $219K/mo total. He wrote the book MAKE about the method: readmake.com.
How he runs it: one person, one loop
He works alone: "I still work alone... I haven't hired at least not for product stuff." One friend keeps the server alive. That is the whole org chart. His loop (levels.io/startups): idea, build, launch, grow, monetize, automate, repeat. "You have an idea, or I would have a problem and make it into an idea. I would build it, I would launch it, I would grow it, and then I would monetize it to make money from it, and then, if I got really annoyed with working on it, I would automate it with robots." Automation is a formal step, not an afterthought: APIs, win-back emails, automatic refunds, automated support. No VC is load-bearing: no board, no consensus, no permission needed to pivot, kill, or launch.
The stack, and what AI changed
Famously boring: PHP, jQuery, SQLite, often a whole product in one index.php file. On the Lex Fridman podcast: "It's all jQuery... It's PHP and jQuery, yes, and SQLite." Then AI coding agents multiplied his speed: nearly a year coding almost solely on his VPS with Claude Code (his writeup), deploys landing in seconds, two ten-second outages in twelve months. August 2026 (his archive): built InfiniteSlop.ai in a single day from his phone, in a sauna, over Termius to a Hetzner VPS running Claude Code. His 2026 punchline (levels.io): "Everyone can now build apps with AI so now distribution is the real challenge." The single most important sentence in the playbook: the build bottleneck is gone, distribution is the game now.
The business playbook
- 12-in-12 origin: 2014, music income collapsing, announced 12 startups in 12 months. Nomad List was #4, RemoteOK #7. The challenge was never about twelve successes; it was about becoming someone who ships.
- Shotgun, not sniper: "Only 4 out of 70+ projects i ever did made money and grew." Plan for a 95% miss rate; make each shot cheap and fast.
- Solve your own problem, start tiny: Nomad List (he needed city data on $700/mo), Hotelist.com (booking sites delete negative reviews). You are the expert on your own problems.
- Launch fast and in public: Nomad List began as a Google spreadsheet with editing accidentally left on. Minimal slice, real payments, public from day one.
- SEO as an asset: Nomad List ranks for thousands of cities because every city is its own data page. Programmatic pages for real demand.
- Neighbor products: RemoteOK was the next thing the same audience needed. Each product feeds the next.
- Tiny products with concrete headlines: pay, upload, get result. One narrow job, money from day one.
- Public revenue as marketing: bio lists revenue per product; X payouts hit ~$25K in Aug 2026. The platform pays him to be there.
- Creators over press: pays niche TikTok creators for demo videos instead of chasing press.
Monetization and pricing: the charge-from-day-one machine
Levels treats pricing as a product decision, not an afterthought. The whole money system runs on one idea: get paid from day one, charge enough per user that tiny numbers make a good business, and automate every cent afterward.
- Subscription where there is repeat, one-time where the result is one-time. That clean split matches what he actually runs: subscriptions on ongoing value (databases, job boards, AI generations), one-time prices on finished results. In his Bali talk he put the compounding math on a chart: a $75 one-time payment growing 25% a year reaches about $183K by year five, while the same product sold as a subscription reaches nearly $2M. Subscriptions are annoying for users, he admits, but the revenue compounds and the math survives churn. The point is not "always subscribe"; it is to pick the model that matches the repeat and understand what you are trading away. Source: levels.io/startups transcript
- Free users never converted, so he stopped building for them. On the Lex Fridman podcast he said the free-user path "never worked for me well" because "free users generally don't convert." The VC version (raise money, buy ads, convert a fraction) only works with a conversion machine he does not have and does not want. His rule: show the landing page and a demo, then "if you want to use it, pay me money." The payment is the validation; there is no free tier to learn from. Source: Lex Fridman #440, cited from our Oct 3 verified notes
- Charge real money: $30-plus per user per month. His line from the same interview: Netflix can charge $10 because Netflix is giant; an indie needs "at least $30 or more on a user to make it worth it." His math is built for small numbers: a thousand users at $30 a month is $30K a month, "and it's a lot of money." The playbook math from earlier still holds: 26 customers at $99 a month is an average U.S. income. Price so a small audience is a good business.
- He has receipts on price ladders. After a decade of charging for Nomads.com, he made it effectively free in September 2026 ($1 to sign up, just to block spam) and published his own paywall-era numbers: at $299 he got 50 new members a month, at $99 he got 200, at $29 he got 400. Total: 43,252 paid members over 12 years, about 300 a month. His two reasons: he had made enough money from it, and charging "artificially limits the size." The next model he names outright: more people plus more sponsors (SafetyWing sponsored the site for years), and the strategic reason: post-AGI, "communities also have a moat where many of my SaaS apps do not." Source: levels.io, opened and read live Oct 5, 2026
- High-ticket sales with zero sales calls. RemoteOK job posts run roughly $100 to $1,000 each, and he sells job-post bundles worth up to $50K through a self-serve Stripe page: the buyer configures the bundle, the discount computes itself, they pay by card. No demos, no discovery calls, no account executive. The pricing page is the sales team. A published summary of his talk; bundle figures are his recounting, not audited
- The money runs on robots. From his own talk transcript: at the time he spoke, 187 parallel processes were running on his server, pulling weather data for Nomad List cities, fetching job posts for RemoteOK, and processing refunds automatically. Automatic refunds are a feature, not a cost: instant refunds plus fast support are how a one-person operation with no face carries trust.
- Margin discipline as a pricing input. Each Interior AI rendering costs him about a cent in compute, and he has reported months at roughly 80% margin (per our Oct 3 notes). The rule underneath: never build unit economics on the assumption that GPU or API prices keep falling; price so you survive if infrastructure gets more expensive.
The money rules, copy-ready:
- One price you understand in five seconds, visible before signup.
- Paid from day one; the payment is the validation.
- $30-plus per user per month unless the model is genuinely one-time.
- Subscription where there is repeat, one-time where the result is one-time.
- No free tier as a growth strategy; a $1 anti-spam wall beats a free tier that teaches you nothing.
- The pricing page is the sales team: self-serve, Stripe, discounts computed in the page.
- Automate refunds and win-backs; trust is the product when there is no face.
- Price ladders are data: test the $299 vs $99 vs $29 curve and watch signup velocity, like he did.
The Studio parallel. Our first-dollar lane is Tugboat's $75 WebFix, and his rules bless it: one simple flat price, understood in five seconds, paid from day one, no free audit tier. The gap: we still hand-draft every outreach email and the offer does not yet live on a page that sells itself, and refunds plus win-backs are not yet automated in the loop. The concrete copy for this week (proposed to Dave): adopt the eight money rules above as standing Studio pricing guidance, and point the WebFix offer at a self-serve page where the pricing page does the selling.
What a non-programmer should copy, ranked by impact
- The bottleneck moved: building is no longer the moat, distribution is.
- Run the shotgun: twelve small bets beat one big one.
- Mine your own problems: lived experience is the niche map.
- Tiny, paid, fast: "You only need 26 customers paying you $99/mo to make an average U.S. income."
- Publish the numbers: revenue transparency works under a pseudonym too.
- Build SEO pages and neighbor products that feed each other.
- Automate at the end of every loop: support, refunds, emails, onboarding.
- No meetings, no employees, no permission: for a non-programmer, the "team" is coding agents on a server.
- Never build the factory before the traffic: ship ugly, get users, then polish.
The stealth adaptation: a faceless Levels-style portfolio
What requires the face: the X audience engine (~966K followers over a decade+). Launches get a day-one crowd money cannot buy cheaply. Build-in-public content, the book as personal artifact, person-profiled press: all lean on visibility. Strip the face and launches start cold.
What does not require a face: almost everything else. The products, programmatic SEO, tiny tools with clear headlines, paid creator demos, affiliates, the whole build loop, automation-first ops. None of it needs Dave's name or photo.
Distribution without a face: SEO pages that earn their own traffic (the Nomad List model, faceless by nature); pay niche creators per video (the creator's face sells, not yours); marketplace and app-store search; acquire small existing products and grow them under their own brands; pseudonymous product accounts publishing revenue numbers; affiliates on a cut.
Trade-offs, stated plainly: a faceless start is slower and costs real money early (creators, affiliates, SEO time). Press is harder without a human story; the counter is radical transparency (public numbers, fast support, real refunds). What you gain: privacy, no brand risk, no content treadmill, products far easier to sell one day. His loop works under a pseudonym. The only part truly lost is the free launch-day crowd, and money plus SEO can buy a slower version of it.
Study the master, adapt the method, drop the face.
Connectors: wiring an orchestrator to external AI services
An agent-first company constantly needs its orchestrator to reach services that do not live inside the agent. Image generation, video generation, payments, publishing: each needs a connector. Three patterns, in order of preference.
Pattern 1: Native API
If the service has an API, use it. Clean auth, structured errors, predictable cost. This is always the first choice.
Pattern 2: MCP servers
The Model Context Protocol is the emerging standard connector: wrap a service once as an MCP server, and any compatible agent can discover and call its tools. Prefer remote servers for reach, group tools around user intent rather than mirroring API endpoints one to one, and defer loading tool definitions the agent does not need yet. New connectors should be built here when possible.
Pattern 3: Browser automation
Some services have no API and no MCP server. The Studio's real example: generating images with ChatGPT's image model and videos with Seedance/Dreamina, both living behind web apps. The pattern is a controlled browser session, signed in once, driven by the agent: navigate, upload, prompt, download the result, verify it with human eyes. The fallback pattern, and it works, but it carries rules:
- Credential hygiene is non-negotiable. Logins live in a secure vault or the platform's approved credential store, never in prompts, logs, memory files, or generated code. Two-factor challenges go to the human's own device.
- Scope the session. The browser profile used for automation is separate from daily browsing. The agent gets the minimum permissions for the named task.
- Verify before using. Anything produced through a browser connector gets checked by a human (or a second agent) before it ships.
- Prefer API the moment one exists. Browser automation is brittle: layouts change, challenges appear, sessions expire. Every browser connector carries a standing note to migrate when a stable interface appears.
- Rate discipline. Drive the service the way a careful human would. A provider rate limit is a hard stop, not a signal to retry harder.
The Studio's connector map today
The orchestrator reaches ChatGPT's image generation through a signed-in browser session for the 4-image review blocks, and reaches Seedance/Dreamina the same way for video generation. Both are Pattern 3 connectors with human verification on every output. Direction of travel: keep them working, migrate either to API or MCP the day a stable interface appears.
Integration checklist for the Studio
What "agent-first" concretely means for this company, mapped to what already exists.
- One orchestrator, many workers. The main agent coordinates; specialized workers run in the background. Keep the hierarchy shallow and the handoffs explicit.
- Every background task is visible. Anything over 30 seconds gets a queue entry with an honest ETA and a real completion check. The human reads the queue, not the logs.
- Human approval on every send, spend, and publish. No exceptions. Outreach, posts, and payments wait for exact-wording sign-off.
- Evals on the critical paths. The queue watchdog is a primitive eval: every item carries a check that proves the work is done. Extend the pattern to content quality and outreach drafts.
- Memory in files, not just in heads. Standing decisions live in the memory system and workspace docs so any worker can pick up the context. Curate it; stale memory is worse than none.
- Connectors documented with their migration path. Each external service gets a note: which pattern it uses today, what would replace it, and who owns the credentials.
- Snapshots so nothing is ever lost. The site is committed and pushed on a schedule. An agent-first company that can lose its own work to a bad deploy is not agent-first.
- Cost tracked per run. Token and service costs get logged next to the business result they produced. If a workflow costs more than the value it creates, redesign the workflow.
- Rebuild without ego. Notion needed five tries. When a harness gets bulky, slow, or expensive, simplify it one component at a time and measure what breaks.
- Ship on a cadence. Daily briefs, hourly shifts, weekly planning. Agents thrive on rhythm; the schedule is the product as much as the output.
← All research