Skip to content
FrankX.AI
Research Hub

Agentic Life Observatory · verified 2026-07-17

Choose infrastructure by what it proves.

A living market map for systems that claim to remember, coordinate, protect, evaluate, or operate across a life. Five failure modes. Six strategic roles. One evidence trail.

Current registry

Systems
29
Categories
9
Coverage
71%
Axes
5

Operating loop

Install. Test. Observe. Evolve.

Every technology enters through the same reversible sequence. No vendor becomes authority because its demo looked good.

  1. 01

    Install

    Connect through a reversible adapter. Name the owner, version, authority boundary, and export path.

    Adapter contract
  2. 02

    Test

    Run recall, deletion, trajectory, policy, and handoff contracts against real tasks before promotion.

    Audit JSON
  3. 03

    Observe

    Compare traces, errors, overrides, cost, latency, and quality. Final answers alone do not pass.

    Trace set
  4. 04

    Evolve

    Promote, constrain, replace, or remove from evidence. Update the registry and preserve the decision.

    Decision record

Market map

Build, integrate, partner, compete, or watch.

Scores run from 0 to 3 across the five structural failure modes. They are directional editorial assessments, not certifications.

29 of 29 systems · 71% coverage

Agentic Architecture Field Guide

FrankX stack

FrankX

Public implementation reference for composing agent runtimes, memory, policy boundaries, evaluation, and domain modules.

Decision detail +

Next: Use the guide as the public architecture baseline and attach executable conformance receipts to each pattern.

Risk: A reference architecture can look complete before its contracts are proven against production workloads.

public reference · self-hosted · See repository

Build

active

Context
Compose
Sovereign
Verify
Life-wide

Agentic Operating System Standard

FrankX stack

FrankX

Portable contracts for modules, agents, skills, workflows, loops, ledgers, gates, adapters, and profiles.

Decision detail +

Next: Stabilize the public conformance profile and adapter contract.

Risk: A standard without conformance tests can become documentation theater.

open-standard · self-hosted · See repository

Build

active

Context
Compose
Sovereign
Verify
Life-wide

Hermes Agent

Harness

Nous Research

Extensible agent harness with tools, skills, memory, scheduling, projects, and local execution surfaces.

Decision detail +

Next: Use as a replaceable control plane while keeping life state in local_core.

Risk: Harness memory and schedules must not silently become the only system of record.

local · self-hosted · See repository

Integrate

active

Context
Compose
Sovereign
Verify
Life-wide

OpenFang

Runtime

RightNow AI

Open-source agent operating system in Rust with local-first runtime ambitions.

Decision detail +

Next: Study runtime packaging and installation; differentiate with life modules, portable memory, and receipts.

Risk: Agent OS positioning is broader than proven multi-domain user outcomes.

local · self-hosted · See repository

Compete

reference

Context
Compose
Sovereign
Verify
Life-wide

LangGraph

Orchestration

LangChain

Graph-based durable orchestration for stateful, human-in-the-loop agent workflows.

Decision detail +

Next: Use for durable graph workflows where replay and interrupts justify framework weight.

Risk: Framework state is not a substitute for portable life memory or cross-harness identity.

cloud · self-hosted · MIT

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Starlight Memory

Memory

FrankX

Sovereign local_core memory authority with replaceable provider adapters and provenance-aware atoms.

Decision detail +

Next: Ship adapter conformance tests for Mem0, Graphiti, and file-native storage.

Risk: Multiple writers require conflict policy and actor-aware provenance.

local · self-hosted · See repository

Build

active

Context
Compose
Sovereign
Verify
Life-wide

Arize Phoenix

Observability

Arize AI

Source-available AI observability and evaluation platform with OpenTelemetry and OpenInference support.

Decision detail +

Next: Prototype local OpenTelemetry traces for agent trajectories and memory operations.

Risk: Observability volume and sensitive spans require retention and redaction controls.

cloud · self-hosted · Elastic License 2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Cognee

Memory

Cognee

Open data and memory engine that turns information into graph-oriented AI memory.

Decision detail +

Next: Compare graph construction, provenance, and deletion semantics with Graphiti.

Risk: Graph ingestion quality and deletion semantics need workload-specific verification.

cloud · self-hosted · Apache-2.0

Compete

monitor

Context
Compose
Sovereign
Verify
Life-wide

Langfuse

Observability

Langfuse

Open-source LLM engineering platform for traces, prompt management, datasets, and evaluations.

Decision detail +

Next: Compare local deployment cost and export fidelity with Phoenix.

Risk: Self-hosting still requires governance for PII, secrets, and retention.

cloud · self-hosted · MIT core with commercial components; verify current repository

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

LangSmith

Observability

LangChain

Managed tracing, evaluation, deployment, and observability tightly integrated with LangChain and LangGraph.

Decision detail +

Next: Use only where LangGraph trace depth outweighs sovereignty and stack coupling.

Risk: Tight ecosystem coupling and hosted traces can reduce portability.

cloud · enterprise · Commercial service; SDKs vary

Watch

monitor

Context
Compose
Sovereign
Verify
Life-wide

Letta

Runtime

Letta

Stateful agent runtime built around self-editing memory blocks and long-lived agents.

Decision detail +

Next: Study memory-block ergonomics and compare against atom-plus-adapter authority.

Risk: Long-lived agent state can couple memory authority too tightly to one runtime.

cloud · self-hosted · Apache-2.0

Inspire

reference

Context
Compose
Sovereign
Verify
Life-wide

MLflow GenAI

Evals

LF AI & Data / Databricks ecosystem

Open platform extending experiment tracking, tracing, evaluation, and model lifecycle management into GenAI applications.

Decision detail +

Next: Evaluate as a durable experiment ledger where existing MLflow infrastructure exists.

Risk: Platform breadth can add operational weight for small local-first teams.

cloud · self-hosted · Apache-2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

n8n

Automation

n8n

Workflow automation platform with self-hosting, integrations, and AI workflow nodes.

Decision detail +

Next: Use as automation fabric for reversible connectors, not as the life-state authority.

Risk: Workflow JSON can become hidden business logic without tests and ownership boundaries.

cloud · self-hosted · Sustainable Use License / Enterprise License

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Starlight Evals

Evals

FrankX

Receipt-first evaluation substrate for private estate tasks, model arenas, and maker-not-equal-checker verification.

Decision detail +

Next: Add shared receipt schema, baseline suites, and CI adapters for all life modules.

Risk: Estate-specific suites can overfit unless public and adversarial tasks remain in the mix.

local · self-hosted · See repository

Build

active

Context
Compose
Sovereign
Verify
Life-wide

AG2

Orchestration

AG2 community

Open multi-agent conversation and workflow framework derived from the AutoGen ecosystem.

Decision detail +

Next: Monitor evolution and compare handoff semantics with LangGraph and ADK.

Risk: Conversational agent loops can be token-heavy and difficult to bound.

self-hosted · Apache-2.0

Watch

monitor

Context
Compose
Sovereign
Verify
Life-wide

Claude Code

Harness

Anthropic

Terminal coding harness with tools, hooks, subagents, skills, MCP, and repository-native instructions.

Decision detail +

Next: Treat as a replaceable execution harness fed by repo contracts and sovereign memory.

Risk: Harness-local context does not compound automatically across other tools.

local-client · cloud-model · Commercial product

Partner

active

Context
Compose
Sovereign
Verify
Life-wide

CrewAI

Orchestration

CrewAI

Role-oriented multi-agent framework and enterprise automation platform.

Decision detail +

Next: Track enterprise control and eval depth; avoid role-play orchestration without evidence contracts.

Risk: Role labels can create multi-agent theater when scopes, stop conditions, and receipts are weak.

cloud · self-hosted · MIT

Watch

monitor

Context
Compose
Sovereign
Verify
Life-wide

Google Agent Development Kit

Orchestration

Google

Open agent framework with multi-agent composition, tools, sessions, evaluation, and deployment integrations.

Decision detail +

Next: Evaluate as an alternate orchestration adapter and A2A reference implementation.

Risk: Cloud deployment convenience can obscure where durable state and traces become authoritative.

cloud · self-hosted · Apache-2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Graphiti

Memory

Zep

Open temporal knowledge graph engine for agent memory and changing entity relationships.

Decision detail +

Next: Test temporal recall where entity history justifies graph complexity.

Risk: Operational complexity may not pay for simple preference or episodic memory.

self-hosted · Apache-2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Model Context Protocol

Protocol

Anthropic / Linux Foundation ecosystem

Open protocol for exposing tools, resources, and prompts to AI applications through stable server contracts.

Decision detail +

Next: Standardize life-module tools and resources without placing authority in any one client.

Risk: Protocol interoperability does not automatically provide identity, policy, memory, or evals.

open-standard · local · cloud · See specification and SDK repositories

Partner

active

Context
Compose
Sovereign
Verify
Life-wide

OpenAI Agents SDK

Orchestration

OpenAI

Provider SDK for agents, handoffs, guardrails, sessions, and tracing.

Decision detail +

Next: Keep behind provider adapters; test handoff and trace export against local receipts.

Risk: Model and hosted-service coupling can weaken portability.

cloud · self-hosted-client · MIT

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

OpenAI Codex

Harness

OpenAI

Coding agent surfaces for repository work, cloud tasks, and terminal execution.

Decision detail +

Next: Keep git and receipt contracts as the coordination plane across Codex and other harnesses.

Risk: Cloud task context and local session context can fragment without shared contracts.

local-client · cloud · Commercial product; CLI licensing varies

Partner

active

Context
Compose
Sovereign
Verify
Life-wide

Braintrust

Evals

Braintrust

Evaluation and observability platform built around datasets, experiments, scorers, and production traces.

Decision detail +

Next: Benchmark CI ergonomics and export receipts against Starlight Evals.

Risk: Hosted eval data needs explicit privacy-class routing.

cloud · hybrid · See SDK repository and service terms

Partner

candidate

Context
Compose
Sovereign
Verify
Life-wide

Cursor

Harness

Anysphere

AI-native code editor with repository context, rules, agents, and background execution.

Decision detail +

Next: Track rules, background-agent receipts, and portability across editor and external harnesses.

Risk: Editor-native memory and rules can become another isolated context island.

desktop-client · cloud-model · Commercial product

Watch

monitor

Context
Compose
Sovereign
Verify
Life-wide

Mem0

Memory

Mem0

General-purpose memory layer for agents with managed and open-source deployment paths.

Decision detail +

Next: Benchmark as a derived adapter against local_core recall and deletion contracts.

Risk: Managed and open-source capabilities differ; provider IDs must not become authoritative.

cloud · self-hosted · Apache-2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Ragas

Evals

Ragas

Open-source evaluation framework focused on retrieval, generation, and agent quality.

Decision detail +

Next: Use for retrieval quality while keeping policy and trajectory checks separate.

Risk: RAG metrics cover only part of a multi-tool agent trajectory.

local · cloud · Apache-2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Supermemory

Memory

Supermemory

Memory and context infrastructure for AI applications with managed developer APIs.

Decision detail +

Next: Track API and self-host maturity; compare recall against Mem0 and local_core.

Risk: Managed context convenience can create another non-portable authority layer.

cloud · self-hosted · See repository

Compete

monitor

Context
Compose
Sovereign
Verify
Life-wide

Agent2Agent Protocol

Protocol

Google / Linux Foundation ecosystem

Interoperability protocol for discovery and collaboration between independently built agents.

Decision detail +

Next: Prototype a bounded A2A handoff with receipts before making it a fleet default.

Risk: Agent interoperability can increase attack surface and authority ambiguity.

open-standard · cloud · self-hosted · Apache-2.0

Partner

candidate

Context
Compose
Sovereign
Verify
Life-wide

DeepEval

Evals

Confident AI

Open-source, test-oriented framework for LLM and agent evaluation.

Decision detail +

Next: Adapt deterministic life-infrastructure contracts into a pytest-style suite.

Risk: LLM judges still require calibration and cannot replace hard policy checks.

local · cloud · Apache-2.0

Integrate

candidate

Context
Compose
Sovereign
Verify
Life-wide

Execution sequence

What to drive next.

Depth before breadth: prove authority, composition, and quality before adding more products.

Now01

Make the substrate testable

  • Conformance tests for memory adapters
  • One receipt schema across life modules
  • Privacy-class routing before cloud writes
Next02

Prove composition

  • One MCP module and one A2A handoff
  • Replayable multi-harness workflow
  • Independent checker on every consequential lane
Scale03

Measure compounding

  • Context rehydration over 30 and 90 days
  • Cross-domain transfer without privacy leakage
  • Replace one vendor without losing authority

Method

A score is a question, not a verdict.

Each score asks whether public evidence shows the system addressing one structural failure mode. A 3 means core to the product. A 0 means not evident. It does not certify security, performance, or fitness for your data.

Refresh claims from primary sources. Test adapters with private workloads. Keep consequential actions human-gated.

Context
Context can persist and improve across tools or sessions.
Compose
Agents, skills, protocols, and workflows can compose through explicit contracts.
Sovereign
Users can export, rehost, audit, or self-host meaningful system state.
Verify
The system exposes traces, tests, evaluations, or machine-checkable receipts.
Life-wide
The system can support more than a single narrow feature or workflow.

Research graph

Read the four foundations behind the map.

Architecture defines the substrate. Memory compounds context. Sovereignty preserves ownership. Evals make quality replayable.

Start with architecture