The AI Agents Worth Evaluating in August 2026
Mira, Viktor, Hermes, ChatGPT Work, and Claude Cowork compared by job, cost, and capability. What to evaluate on August 21, 2026.
Leave with a one-week evaluation plan and a shortlist of agents matched to the job you actually have.
TL;DR: Stop ranking AI agents as if they were one product. On 21 August 2026 the useful split is by job, not by model brand. Mira is a Telegram-native personal agent (free, then Pro from $27/mo). Viktor is a Slack/Teams coworker with a cloud computer (free $100 credits, paid from $50/mo). Hermes Agent is a local, MIT-licensed runtime you can point at 300+ models. ChatGPT Work (9 July 2026) and Claude Cowork (computer use generally available 20 August 2026) closed a lot of the "chatbots only answer" gap. Evaluate those five first. Add OpenClaw, Manus, and a coding harness only if that is the actual work.
I keep getting the same pitch. A new agent "kills ChatGPT." Another one "replaces Claude." Someone in Telegram forwards mira.tg and asks why I still run Hermes. Someone in Slack forwards Viktor and asks why I still pay for Claude Max.
The honest answer is boring, which is why it does not travel well on X: those products do not compete for the same hour of your day. A messenger agent, a Slack employee, a local runtime, and a flagship chat app with computer use are four different machines. If you force them onto one leaderboard, the marketing department always wins and your calendar still looks the same.
This is the field guide I wanted before spending another week in onboarding screens. Prices, docs, and product names were checked against primary pages on 21 August 2026. If a number moved after that date, trust the vendor page over this post.
Which AI agents are actually worth evaluating in August 2026?
Evaluate these nine. Skip the rest until one of these fails a real task.
| Agent | Surface | Job it owns | Start cost (21 Aug 2026) | Primary docs |
|---|---|---|---|---|
| Mira | Telegram (Messages/Slack marked "soon") | Personal and group chat that takes actions | Free, then Pro from $27/mo, Pro Max $99/mo (Mira vs Viktor) | mira.tg, H1 2026 report |
| Viktor | Slack, Microsoft Teams | Shared team coworker that ships files, apps, reports | $100 free credits, paid from $50/mo (Viktor pricing; Mira's compare page also lists Team at $100/mo) | viktor.com |
| Hermes Agent | CLI, desktop, Telegram/Discord/Slack and 20+ gateways | Sovereign runtime: skills, memory, cron, git, tools | Software $0 (MIT). Inference via your keys or Nous Portal Plus $20 / Super $100 / Ultra $200 | Hermes docs |
| ChatGPT Work | ChatGPT web, mobile, desktop | Multi-hour knowledge work that returns docs, sheets, slides, sites | Included with ChatGPT plans. Plus $20, Pro $200. Work started on Pro/Enterprise/Edu on 9 July 2026 (OpenAI) | Same |
| Claude Cowork | Claude desktop, web, mobile, Chrome side panel | Knowledge work plus computer use on your machine | Paid Claude only (Pro $20, Max $100 or $200) (Anthropic Help Center) | Cowork safety |
| OpenClaw | Local gateway into WhatsApp/Telegram/Slack/Discord | Messaging-first open agent you host | Software $0. You bring model spend. Security overhead is real | OpenClaw explainer |
| Manus | Hosted cloud computer | Broad research, sites, decks, spreadsheets you do not want on your laptop | Free/credits vary; paid from about $20/mo for ~4,000 credits (Manus help) | manus.im/pricing |
| Claude Code / Codex / Cursor / Grok Build | Terminal or IDE | Software engineering | $20 to $200/mo depending on harness | Coding-agent guide |
| Antigravity / Gemini CLI path | Google's coding CLI | Google-ecosystem coding agents | Consumer Gemini CLI was moved toward Antigravity CLI in June 2026 | Google developer posts |
If you already compared ChatGPT, Claude, and Gemini as $20 chat apps, that write-up still holds for everyday Q&A: ChatGPT vs Claude vs Gemini 2026. For the standing landscape view, use the Agent Hub. This post is the 21 August 2026 evaluation pass: what to try this week, what it costs, and which job each agent actually owns.
What job does each agent own?
Four jobs. Most "this kills Claude" threads collapse at least two of them.
1. Personal operator in a messenger. You already live in Telegram (or will). You want memory, group summaries, a reminder, an image, a calendar hold, without installing another IDE. That is Mira. OpenClaw and Hermes can sit in Telegram too, but they ask you to run a process.
2. Shared coworker in the company chat. The work is already in Slack or Teams. The unit of work is a thread, a report, a live URL, a CRM update. That is Viktor. ChatGPT Work and Claude Cowork can connect to Slack. They are not Slack-native employees.
3. Local runtime you own. You want skills that accumulate, cron that fires while you sleep, git worktrees, approval gates, and the right to swap GPT-5.6 for Opus 5 for Grok for a local model mid-week. That is Hermes. OpenClaw is the other serious option in this lane.
4. Flagship chat that now finishes the artifact. You want the model lab's own agent, in the app you already pay for, returning a deck or a sheet or a clicked-through browser session. That is ChatGPT Work and Claude Cowork. They got much closer to "does the work" in July and August 2026. The challenger marketing has not fully caught up.
Coding is a fifth job. I will not re-rank Claude Code against Cursor here. I already did that in the 2026 coding-agent guide and the Claude Code pricing note. Treat those as the engineering shortlist, not as personal assistants.
Is Mira better than ChatGPT or Claude?
No. Mira is better at living in Telegram. ChatGPT and Claude are better at being the model lab's own product. If you flatten that into a single "smarter" ranking, you will buy the wrong thing.
Mira's homepage (checked 21 August 2026) sells a personal agent you "hire" inside Telegram: content, inbox, crypto, shopping, travel, group-chat teammate, 1,000+ tools, zero setup, no credit card. It also prints 1,000,000 people hired Mira. Its own H1 2026 report is more useful than the hero stat, because it publishes behavior rather than slogans:
- 550K monthly active users in H1, plus 175K group chats.
- 460 billion tokens processed on OpenRouter in 2026, enough for Mira to call itself a top-3 personal AI agent on OpenRouter.
- 310K images, videos, and tracks generated in H1 (81% images, 13% video, 6% music).
- 29 million+ tool actions in H1, about 23 per active user. One in eight users chained 3+ tools in one task.
- 40K+ people built at least one skill.
- Usage beyond chat: live web search 29%, images 8%, video 3%, reminders 2%, connecting Gmail/Calendar/Sheets 0.4%.
Read that last line twice. Mira's marketing says "other AI tools tell you what to do, Mira actually does things." Mira's own telemetry says most people still chat and search. Connecting the external services that make an agent an operator is still a 1-in-250 behavior. That is not a dunk. It is the same adoption gap every agent product has. If you evaluate Mira, the test is not "can it talk." The test is whether you will connect Gmail and Calendar and then trust it to act.
Two more facts from the same report, because they kill a different myth:
- No frontier closed model sits in Mira's top five. Across those 460B OpenRouter tokens, roughly 95% of volume ran on open-weight models. The named mix: MiniMax M3 37%, gpt-oss-120b 16%, GLM 4.7 Flash 15%, MiniMax M2.5 10%, MiMo-V2-Flash 7%. Mira can route to Claude, GPT, Gemini, Grok, Kimi on paid plans. At population scale it does not.
- Voice is the power-user tell. Mira says 32% of users send voice, and those users almost do not go back to typing. If you will not talk to a bot in Telegram, you are not Mira's core user.
Privacy claims on the homepage: data not sold or used for training, SOC 2 Type I, Tier 3 CASA, GDPR, CCPA Ready, plus a Private Mode that Mira describes as confidential compute. I have not audited those controls. Treat them as vendor claims and read the privacy policy before you paste a customer list into a group chat.
Evaluate Mira this week if: Telegram is already your home screen, you want group-chat memory, and you will connect at least one real tool (Gmail, Calendar, GitHub). Skip it if your work is a git repo or a Slack workspace. Mira is not Claude Code, and it is not Viktor.
What is Viktor, and who should hire it?
Viktor's own line is "not a tool, a hire." The product is a Slack- and Teams-native coworker with its own cloud computer. It writes and runs code, ships a live URL instead of a gist, connects to 3,200+ tools, and invents an integration from docs when one is missing (viktor.com; Mira's compare page is actually a fair summary of the split).
That is a different bet from Mira's. Viktor assumes a company workspace, approval gates, per-person permissions, and SOC 2. Mira assumes you and your groups. If your team does not live in Slack or Teams, Viktor has nowhere to stand.
Pricing, as published and cross-checked on 21 August 2026:
- Start: $100 in credits, no card. Mira's compare page says those free credits never expire.
- Public paid: from $50/month for the workspace, not per seat, with 20,000 monthly credits on the Team plan in several independent write-ups (Viktor pricing, eesel on the credit model).
- Mira's compare page lists Team at $100/mo with 40,000 shared credits that roll over. Vendor pages have moved this year. Confirm the live card before you forecast a bill.
- Credits map to model usage (Anthropic, OpenAI, and others). Quick tasks are quoted around 100-300 credits. Full projects can run to 5,000. That is why a team doing real volume can land well above the sticker. One August 2026 review put heavy use around 9,000 credits/day and a $500-$750 monthly bill (Efficient App).
Independent reviews keep repeating the same split, and I agree with it: Viktor executes, ChatGPT/Claude explain, Cursor generates code, Viktor returns a running thing. The cost of that execution is metering. If your work is bursty and high-token, the $50 headline is a door price, not a ceiling.
Evaluate Viktor this week if: the team already pays for Slack or Teams, you have a recurring report / CRM / ops loop, and you will assign one human owner for approvals. Skip it if you need a personal Telegram agent, a local git runtime, or a model-lab writing partner. Viktor is a coworker. It is not a replacement for Opus 5 on a 200-page draft.
How does Hermes Agent compare to ChatGPT, Claude, Mira, and Viktor?
I run Hermes every day. That is a conflict of interest, so here is the non-marketing version.
Hermes Agent is open-source software from Nous Research (MIT license, no seat fee). It is in the same category as Claude Code, Codex, and OpenClaw: a tool-using agent with a terminal, files, browser, memory, skills, cron, and a messaging gateway. It is not in the same category as ChatGPT Plus. ChatGPT is a product. Hermes is a runtime you point at products.
What that means in practice:
- Any model. Nous Portal, OpenAI, Anthropic, Google, xAI, DeepSeek, Kimi, GLM, MiniMax, local Ollama, Copilot, Codex OAuth. Swap mid-workflow. Mira and Viktor also route across models. They do not let you run the loop on your disk with your
AGENTS.md. - Any surface. CLI, Ink TUI, native desktop, web dashboard, ACP for IDEs, plus Telegram, Discord, Slack, WhatsApp, Signal, iMessage, Matrix, Teams, email, and more. Mira is Telegram-first (other messengers "soon"). Viktor is Slack/Teams-first. Hermes is the same agent behind all of those pipes.
- Memory and skills that you own. Cross-session memory, skill files that accumulate when a workflow works, a curator that archives stale skills instead of deleting them. This is the part ChatGPT memory still does not replace, because the memory lives next to your tools.
- Cron, webhooks, kanban, subagents. Recurring jobs, inbound webhooks, a durable work queue,
delegate_taskfor parallel workers. That is closer to an ops layer than to a chatbot. - Cost is three layers. The binary is free. Inference is your API bill or a Nous Portal plan (Plus $20 with $22 credits, Super $100 with $110, Ultra $200 with $220, checked against public Portal write-ups dated August 2026). Hosted tools (search, images, browser) bill against those credits. Optional Hermes Cloud instances run about $0.29 to $1.09 per day plus inference (Portal info). You can also attach a ChatGPT, Claude Max, or SuperGrok subscription through OAuth and spend that instead of a second API key.
The trade is setup and judgment. Mira's pitch is "10 seconds, no API keys." Hermes will ask you to install a runtime, pick a model, and decide what the agent is allowed to run. If you will not do that, do not evaluate Hermes this week. Evaluate Mira or ChatGPT Work. If you will do that, Hermes is the only agent on this list that can sit on a Yoga Book, receive Telegram, run git, keep skills, and still call Claude, GPT, or Grok without a vendor lock on the loop.
OpenClaw occupies the same local-gateway niche. I covered the Claw family separately: OpenClaw, NemoClaw, NanoClaw. Short version: OpenClaw is messaging-first and exploded in stars. Hermes is Python, skill-centric, and provider-agnostic with a desktop app. Pick one local runtime. Running both as daily drivers is how you lose weekends.
What did ChatGPT Work and Claude Cowork change this summer?
This is the part the "ChatGPT is dead" posts keep missing. The labs shipped agents.
ChatGPT Work launched 9 July 2026 (OpenAI). It is an agent inside ChatGPT, powered by Codex and GPT-5.6, that is supposed to stay on a project for hours and return finished materials: sheets, slides, docs, web apps. Plugins to Slack, Teams, Drive, SharePoint. Scheduled tasks. A desktop app that merged Codex into ChatGPT. Reuters noted the explicit frame: OpenAI going after Claude Cowork and Microsoft Copilot Cowork for non-coders who still need agentic execution (Reuters, 9 July 2026).
If you already pay for ChatGPT Plus or Pro, Work is the first evaluation, not Mira. You already have the account. Give it a task you can grade: a board deck from files you know, a sheet from a CSV you can checksum, a site you can click. If it finishes, you did not need a new vendor. If it stalls, you now have a baseline for Viktor or Hermes.
Claude Cowork is Anthropic's equivalent, on paid plans, across desktop, web, and mobile (Help Center, updated 18 August 2026). Two August updates matter:
- 12 August 2026: the Claude in Chrome side panel became a Cowork session. Skills and connectors work in the tab. A task can start in Chrome and finish on desktop (Anthropic).
- 20 August 2026: computer use, the Skills API, and the Files API went generally available on the Claude Platform. Computer use now takes several GUI actions per turn instead of one round trip, and a browser-use tool was added for page structure rather than pixels only (Anthropic).
Computer use is still a research preview for Pro and Max in the consumer apps, and it is the most dangerous feature on this list. Anthropic says it plainly: unlike sandboxed code execution, computer use has no sandbox between Claude and the screen. Block banking and health apps. Watch the first ten sessions. I would not turn this on for an unattended cron.
Model context, if you are choosing a lab rather than a messenger: Claude's 2026 lineup (Opus 5, Sonnet 5, and the rest) and GPT-5.6 (Sol / Terra / Luna) are covered in the frontier model landscape. API ballpark as of this month: Opus 5 around $5 / $25 per million tokens, Sonnet 5 $2 / $10 (now permanent, Anthropic, 10 August 2026), GPT-5.6 Sol around $5 / $30. Those numbers are why a "free Telegram agent" can still be expensive in aggregate (Mira's 460B tokens), and why a $20 chat plan can still rate-limit you once the agent starts clicking.
How much do these agents cost in August 2026?
Sticker price is a trap. Match the cost shape to the work.
| Cost shape | Who uses it | What goes right | What goes wrong |
|---|---|---|---|
| Flat messenger sub | Mira Pro $27, Pro Max $99 | Predictable. Groups included, no per-seat | You will still hit usage limits. Confirm what "flat" excludes (image/video, premium models) |
| Workspace credits | Viktor from $50/mo + credit burn | Cheap to try ($100 free). Whole team shares | Heavy days turn into $500+ months. Credits are not hours |
| Subscription + included usage | ChatGPT Plus $20 / Pro $200; Claude Pro $20 / Max $100-$200 | One bill, Work or Cowork included | Agent sessions eat the cap faster than chat. Coding and Cowork share the pool |
| Runtime + inference | Hermes $0 + Portal $20/$100/$200 or BYO keys | You see every dollar. Swap models | You own setup, keys, and failure modes |
| Hosted credits | Manus from ~$20/mo for ~4,000 credits | No laptop risk. Good for one-off research/sites | One deep-research run can spend 15-25% of the month (Compass review, Aug 2026) |
| IDE / terminal seats | Cursor Pro $20; Claude Max; Codex Plus $20 / Pro from $100 | Best coding loop | Do not buy this to run your inbox |
A sane stack for a solo operator who already writes and ships:
- One flagship chat plan (Claude Max or ChatGPT Pro, not both, until you can prove the second pays for itself). Use Cowork or Work as the default knowledge-work agent.
- One coding harness (Claude Code, Codex, or Cursor). See the coding guide.
- One runtime if you need Telegram/cron/git on a machine you own (Hermes or OpenClaw, not both).
- Mira only if Telegram groups are a real workflow, not a demo.
- Viktor only if a Slack/Teams workspace will share it.
A team of eight in Slack should invert that: Viktor first, then one lab subscription for writing, then a coding harness for the engineers. Mira is a side channel, not the system of record.
Which other agents still deserve a look?
Manus. Hosted generalist. Cloud computer, research, sites, decks. Meta's acquisition attempt was blocked in April 2026; Manus has been signaling a return to independence on its pricing page. Evaluate it for work you do not want on your laptop. Do not evaluate it as a Slack employee or a git agent.
OpenClaw. Local messaging gateway. Fastest open-source breakout in this category earlier in 2026, and a real security story (exposed instances, malicious skills). If you want this shape with more enterprise wrapping, look at NemoClaw. Details in the OpenClaw piece.
Grok Build. xAI's terminal coding agent. SuperGrok about $30/mo, SuperGrok Heavy listed much higher. Relevant if you already live in Grok. Not a Mira competitor.
Antigravity CLI. Google's direction of travel after consumer Gemini CLI access changed in June 2026. Relevant if your identity is already Google AI Pro/Ultra. Not a personal operator.
Perplexity, ChatGPT wrapper bots in Telegram, image-only bots. Fine as features. They are not agents. Mira's own comparison post is correct on that narrow point (Best AI assistant for Telegram).
If you still need the model-only view, use What is agentic AI? for the definition and the frontier landscape post for benchmarks. Definitions without a job still will not pick a vendor.
How should you evaluate an AI agent in one week?
Do not run a beauty contest. Run four tasks you already know the answer to. Grade artifacts, not vibes.
Day 1. Pick the surface, not the model. If the work arrives in Telegram, Mira or Hermes. Slack/Teams, Viktor. Browser and docs, ChatGPT Work or Claude Cowork. Git, a coding harness. If you cannot name the surface, you are shopping.
Day 2. One connected tool. Gmail, Calendar, Drive, GitHub, or Linear. Mira's own data says most users never do this. If you will not connect a tool, you are buying a chatbot with extra UI.
Day 3. One scheduled job. A morning brief, a weekly report, a watchdog. Cron is the difference between a demo and an employee. Hermes and Viktor are built for this. ChatGPT Work has scheduled tasks. Mira can do it; usage is still ~2% of tasks.
Day 4. One artifact you can checksum. A sheet with known totals. A PR with tests. A live URL. A 1,200-word brief against a source PDF. If the agent cannot produce a thing you can grade, it is still a chatbot.
Day 5. One permission scare. Revoke a tool. Watch what happens. Read the computer-use warning on Claude. Check where Mira stores memory. Check Viktor's approval gates. If you cannot explain the blast radius, do not leave it on overnight.
Day 6. Cost replay. Export usage. For credit products, convert credits to dollars. For subscriptions, count rate-limit hits. For Hermes, read the session cost. If you cannot reconstruct the bill, you cannot keep the agent.
Day 7. Kill or keep. Keep only if it won a task you already pay a human or a Saturday for. Otherwise delete the integration. Stack-creep is how a $20 experiment becomes $400 of overlapping subscriptions.
I will not tell you Mira is a toy or that Hermes makes ChatGPT obsolete. Mira is the right evaluation if Telegram is the work. Viktor is the right evaluation if Slack is the work. ChatGPT Work and Claude Cowork are the right evaluation if you already pay the labs and have not used the agent mode they shipped this summer. Hermes is the right evaluation if you want the loop on your machine. Everything else is a demo until it survives day 4.
FAQ
Is Mira better than ChatGPT in 2026?
For Telegram-native personal and group work, Mira is the more complete messenger agent: memory, groups, skills, 1,000+ tools, free to start (mira.tg). For writing quality, image generation, voice, and the Work agent that returns decks and sheets, ChatGPT is still the broader product (OpenAI on Work). They are not substitutes.
Is Viktor better than Claude?
Viktor is stronger at cross-tool execution inside Slack or Teams. Claude (Cowork + Claude Code + Opus 5/Sonnet 5) is stronger at writing, coding, and computer use on your desktop. Independent roundups keep landing on that split (Vellum on Viktor alternatives, Efficient App). Hire Viktor for ops. Keep Claude for thought and code.
Should I replace Hermes Agent with Mira?
Only if you never needed a local runtime. Mira wins on zero setup and Telegram distribution (550K MAU in H1, Mira report). Hermes wins on owning skills, cron, git, and model choice (Hermes docs). I run Hermes as the control plane and treat messenger agents as edges. If you want one app and no terminal, run Mira and skip Hermes.
How much does a serious agent stack cost per month?
A solo operator can stay near $20-$50 (one Plus/Pro plan, or Mira Pro, or Hermes Plus). A daily driver with Max/Pro, a coding harness, and a credit-based coworker lands at $150-$400. A Slack team that actually uses Viktor can see $500+ when credits burn. Recompute from your own usage logs after a week. Sticker prices on 21 August 2026 are in the table above.
Did ChatGPT and Claude actually catch up to "real agents"?
They closed the obvious gap. ChatGPT Work (9 July 2026) and Claude Cowork plus computer-use GA (20 August 2026) both ship artifacts and take actions, not only paragraphs. They still run inside the lab's app, with the lab's limits, on the lab's memory model. Messenger and Slack agents still win on where the conversation already happens. Local runtimes still win on control.
What is the one-week test I should run?
Connect one tool, schedule one job, demand one checksummable artifact, then replay the bill. Keep the agent only if it wins a task you already pay for. That protocol is in the section above.
Where should I go next on frankx.ai?
Start with the Agent Hub for the standing map, What is agentic AI? if the vocabulary is new, the coding-agent guide if the job is software, and the ChatGPT vs Claude vs Gemini piece if you still need a $20 chat plan. Then subscribe to the newsletter if you want the next pricing and capability pass when these pages move again.
If you want a personal operator stack rather than another chat tab, that is the work I do with ACOS and a Hermes-centered runtime. The short version: one surface for talk, one harness for code, one runtime you own. Everything else is a vendor demo until it survives a week of real tasks.
Build your first AI system
Step-by-step guide to setting up ACOS, creating your first agent, and shipping real products with AI.
Start buildingProduction-ready architecture
Download AI architecture templates, multi-agent blueprints, and prompt engineering patterns.
Browse templatesJoin the builder community
Connect with creators and architects shipping AI products. Weekly office hours, shared resources, direct access.
Join the circleRead on FrankX.AI — AI Architecture, Music & Creator Intelligence
Stay in the intelligence loop
Weekly field notes on AI systems, production patterns, and builder strategy.
Continue Reading

Anthropic Paused the Claude Agent SDK Credit Change. Here's What Builders Sho...
Anthropic paused the Claude Agent SDK credit change. What it means for claude -p, OpenClaw, OpenCode, Codex, and agent pricing.
Read article
OpenClaw, NemoClaw, NanoClaw: The AI Agent Ecosystem
The open-source AI agent with 247K GitHub stars. How it works, what NVIDIA added with NemoClaw, and what it means for personal AI agents in 2026.
Read article
Your Mind Is a Temporary Library
A Mindvalley field note on turning lived experience into durable knowledge, agentic systems, and public capability that can outlast us.
Read article