Which model for which architecture —with sources, tests, and a named author.
The swarm recommends a job and a system shape first, then a model. Vendor scores, independent composites, and first-party receipts stay labeled. Frank still publishes.
104
Research domains
104
Domains with sources
653+
Source references
5
Specialist agent roles
Swarm recommendation · as of 2026-08-16
Architecture first. Then a model.
These are routing cards for how a system should be built, not slogans about FrankX. Each row names the job, the architecture, a primary, an alternate, and the evidence kind. No SIS battle score is implied.
| Job | Architecture | Primary | Evidence | Read |
|---|---|---|---|---|
| Long-running coding and knowledge agents | One reasoning model with tools, verification loops, and a human publish gate. Do not swap the model mid-task. Do not: Do not treat Artificial Analysis Index ties as a “win.” | Grok 4.6 Alt: Claude Opus / Fable 5 when the task is long-horizon and cost is secondary | Independent composite AA Intelligence Index 61 for Grok 4.6 (vendor + AA, 12 Aug 2026). No SIS arena receipt for Grok. | Open |
| Receipt-gated model battles | Dispatch only models the Starlight arena harness can pin. Publish JSON or do not publish a winner. Do not: Do not invent a Grok 4.6 arena winner from vendor benches. | Claude Fable / Opus / Sonnet / Haiku Alt: Hold Grok, GPT, and Gemini until the harness can write a receipt | First-party SIS tools/arena is Claude Code Agent-native. Existing /research/model-arena cards are those receipts. | Open |
| Quiet product still-life | Use a native image tool already in the session. Inspect the pixels. Show only frames that pass QA. Do not: Do not route this work through FAL. Do not publish missed briefs. | Grok Imagine, Codex image_gen, or Antigravity generate_image Alt: Pick the CLI you are already in | First-party Same prompt, 16 Aug 2026. All three returned a publishable still-life. Quality id still did not appear in the Grok tool result. Video not run. | Open |
| Catalog, price, and compare closed models | Keep a registry with live pricing and compare pages. Routing rows must name evidence kind. Do not: Do not dump vendor tables onto Arcanea or starlightintelligence.org. | LLM Hub / AI Ops models catalog Alt: Dated analysis posts when a flagship ships | Vendor-claim Hub registry plus /llm-hub/compare/grok-4-6-vs-grok-4-3. Prices are vendor list. | Open |
| Lowest-cost closed frontier chat | Prefer the cheaper prior Grok when 500k context and agent stamina are not required. Do not: Do not retire 4.3 solely because 4.6 exists. | Grok 4.3 Alt: Grok 4.6 when the task is long-running | Vendor-claim Existing hub routing: 4.3 remains the cheap closed-frontier row. | Open |
Flagship models
Architecture and job, not a crown. Evidence kind is on every card. Registry notes that were not re-run this week stay vendor-claim.
ga · Independent composite
Grok 4.6
Long-running agents at mid price
Post-training / agent RL refresh. 500k context.
AA Index 61 (vendor + AA, 12 Aug 2026). No SIS arena receipt.
ga · First-party
Claude Fable 5
Receipt-gated arena + long-horizon Claude work
1M context. SIS arena can pin Fable. Use when the harness must write a JSON receipt.
Model Arena cards are Claude Code Agent receipts. Launch benches remain vendor-claim.
ga · Vendor-claim
Claude Opus 4.8
Cost-secondary coding and knowledge work
1M context. Point release after 4.7. Registry notes $5/$25 unchanged.
From the in-repo model registry. Not re-benchmarked in this wave.
ga · Vendor-claim
GPT-5.5
Long-context OpenAI agents
1M context. Registry: agentic/long-context step; narrow terminal-agent edge claimed by vendor.
No first-party SIS receipt in this wave.
ga · Vendor-claim
Gemini 3.5 Flash
Shipped Google 3.5 line
1M context. Flash is the GA 3.5 model. Pro remains preview in the registry.
Antigravity generate_image is a separate native image path, not this text model.
previous · Vendor-claim
Grok 4.3
Cheaper closed-frontier chat
1M context. Keep when 4.6 agent stamina is not required.
Previous xAI flagship. Do not retire it only because 4.6 exists.
Executed tests
Only runs that produced an artifact are listed. Holds stay holds.
First-party · 16 Aug 2026
One still-life, three native generators
Grok Imagine, Codex image_gen, and Antigravity generate_image. Only QA-pass frames. No FAL.
SIS receipts · Claude-native
Model Arena
Receipt-gated battles only. No Grok 4.6 card until the harness can pin it.
Vendor + AA · 12 Aug 2026
Grok 4.6 brief
Same-scale post-training refresh. AA Index 61. Labeled scores. Not an arena winner.
Who did what
Model, role, skill, and date for this wave. Human publish still sits last.
01 · 2026-08-14 → 2026-08-16
Grok 4.6 · Hermes default / authoring
Wrote the sourced Grok 4.6 brief, hub registry row, 4.6 vs 4.3 compare page, and this recommendation board.
Skills / tools: content-swarm-production, seo-geo-public-contract-audit, estate-design-excellence
02 · 2026-08-14
Claude Code · Independent read-only review
Confirmed: do not add n8n or Railway MCP for research publishing; Model Arena stays receipt-gated; do not invent a Grok winner.
Skills / tools: claude-code -p, print mode
03 · 2026-08-15
Grok Imagine · Image backend (not Grok 4.6 text)
Returned a curated still-life. Runtime model grok-imagine-image. Quality id still not in the tool result. Video not run.
Skills / tools: hermes image_generate, image-workflow-orchestrator
04 · 2026-08-16
Codex + Antigravity · Native image peers
Same still-life brief. Codex image_gen and agy generate_image both wrote real files that passed QA.
Skills / tools: codex exec image_gen.imagegen · agy generate_image
05 · 2026-08-17
Frank Riemer · Human publish gate
Human publish gate. This wave ships only after PR 483 merges to main and Vercel deploys.
Skills / tools: GitHub PR + Vercel git integration
Questions
- How does the FrankX swarm pick a model?
- By job and architecture, not by a single leaderboard. A recommendation names the primary model, an alternate, what not to do, and the evidence kind: first-party, vendor-claim, independent composite, or not-run.
- Did you run Grok 4.6 against Claude in Model Arena?
- No. Model Arena only publishes SIS JSON receipts. The current harness dispatches Claude Code Agent models. Vendor and Artificial Analysis scores for Grok 4.6 stay labeled and live on /llm-hub/grok-4-6.
- Is Grok Imagine the same as Grok 4.6?
- No. Grok 4.6 is the reasoning model. Still-lifes came from Grok Imagine, Codex image_gen, and Antigravity generate_image.
- Who wrote the Grok 4.6 brief and this hub update?
- Hermes default on Grok 4.6 drafted the packet and this board. Claude Code ran a read-only review of the publish path (no n8n, arena stays receipt-gated). Frank remains the human publish gate.
- Which flagship should I start with?
- Match the job. Long-running mid-price agents: Grok 4.6. Receipt-gated battles: Claude Fable or Opus via Model Arena. Cheap closed chat: Grok 4.3. Google 3.5 that shipped: Gemini 3.5 Flash. GPT-5.5 stays vendor-claim until we run it.
Flagship Articles
Long-form investigations that preserve sources, questions, and the distinction between reported evidence and my interpretation.
New foundational program
Four personal qualities. Four research lenses.
Freedom, Mastery, Meaning, and Connection begin as autobiographical commitments. This research program asks what autonomy, expertise, meaning, belonging, and collective intelligence can responsibly add — without turning a personal constitution into a universal personality theory.
Freedom
Autonomy lens
Mastery
Expertise lens
Meaning
Coherence lens
Connection
Belonging lens
Recently refreshed
The domains I have most recently revisited, with source dates and unresolved questions kept close to the synthesis.
All research domains
104 research areas organized by topic. Specialist agents map the evidence and contradictions; I review what the page can responsibly conclude.
Specialist research roles
Five bounded roles support scanning, evidence review, synthesis, and production. They work inside directed sessions; none has authority to publish on its own.
Frontier Intelligence Analyst
Technology & Market Research
Tracking cutting-edge AI developments, framework releases, and market shifts across the global AI landscape
Systems Architecture Researcher
Infrastructure & Patterns Analysis
Evaluating production architectures, deployment patterns, and infrastructure decisions for enterprise AI systems
Evidence Synthesis Engine
Claims Validation & Cross-Reference
Validating research claims against primary sources, cross-referencing across publications, and maintaining confidence ratings
Strategic Pattern Analyst
Trend Detection & Forecasting
Identifying convergence patterns across domains, detecting emerging trends, and mapping technology trajectories
Publication & Distribution Architect
Content Strategy & SEO/AEO
Transforming validated research into SEO-optimized briefs, AI-citable summaries, and structured knowledge artifacts
Research methodology
Primary sources are preferred. High-confidence claims require independent support; uncertainty stays visible when the material cannot justify certainty.
Directed scan
A defined question guides the search across primary material, research papers, official releases, and credible expert analysis
Specialist passes
Separate agents compare sources, inspect contradictions, and distinguish reported fact from interpretation
Evidence review
Claims are checked against the available sources, given a confidence level, and narrowed when the evidence is incomplete
Human decision
Frank reviews the synthesis, changes or rejects weak claims, and decides whether the artifact is useful enough to publish
From research to practice
Learn the tools hands-on
The research maps the landscape. These portals curate the videos, docs, and expert channels to actually build with each platform.
Claude & Anthropic Mastery
Master Anthropic's full Claude stack — Opus 4.8, Sonnet 4.6, Haiku 4.5, Claude Code, the Agent SDK, MCP, Computer Use, and Skills — from first prompt to production agents.
Codex & OpenAI Agent Mastery
Master OpenAI Codex for agentic software work: setup, local CLI workflows, AGENTS.md, code review, and production-ready iteration.
ChatGPT & OpenAI Mastery
Master ChatGPT for everyday work, prompting, data analysis, custom workflows, and practical OpenAI fluency.
Gemini & Google AI Mastery
Master Google's full AI stack — Gemini 3.5 Flash, Gemini 3.1 Pro, Antigravity 2.0, NotebookLM, Veo 3.1, and Nano Banana Pro — from your first prompt to production agents.
Antigravity Mastery
Master Google Antigravity — the standalone agent-first development platform (desktop app, CLI, SDK) that replaced Gemini CLI — from first install to production multi-agent workflows.
Stay Current
Get weekly intelligence briefs synthesizing the most important developments across AI architecture, production patterns, and emerging technology.