Skip to content

Agent Runtime, Hosting, and Supplier Landscape ​

Working research note v0.3 (updated September 24, 2026; first published September 4, 2026). This is a point-in-time technical survey for the continuing agentbox exploration, informed by the earlier aClaw experiment. It maps options for a future runtime proof of concept; it is not a product announcement, procurement recommendation, or committed roadmap.

The earlier aClaw experiment offered each user an independent OpenClaw or Hermes agent in an isolated container. That architecture was useful because it made the agent concrete: one user, one process environment, one filesystem, one place to work.

It also coupled too many layers together.

In 2026, a stronger way to frame the problem is to separate the system into replaceable parts:

text
Model
  -> Agent harness
  -> Durable identity and state
  -> Private computer or sandbox
  -> Hosting substrate

Each layer is changing at a different speed. Models improve quickly. Harnesses compete on orchestration and memory. Sandbox providers compete on isolation, persistence, wake time, and operational controls. A durable product should not confuse any one of those suppliers with the agent itself.

What changed in v0.3 ​

Meta's Muse adds a consumer-product reference for a personal agent with a dedicated cloud computer. Its separate Sentinel controls connector actions and network egress from outside the main agent's runtime. The noteworthy design is the enforceable boundary around the agent, not simply the presence of a second model. Muse is a complete Meta-operated product, not a hosting supplier agentbox can substitute into a proof of concept. Our essay on Muse discusses its product implications in more detail.

The current agentbox-cloud proof of concept also makes one product form less hypothetical: Hermes running in a Cloudflare Sandbox, with a Worker, Durable Object, D1, and R2 as its control plane. That experiment is distinct from the Cloudflare Agents framework discussed below. Its current security and recovery gaps should be part of the comparison, not hidden by the general strengths of its substrate.

What changed in v0.2 ​

OpenAI's new Agents API changes the center of this comparison. The earlier version treated the OpenAI Agents SDK as a thin harness that agentbox would have to operate. The Agents API now offers a managed Codex harness, durable sessions, context compaction, recovery, subagents, skills, MCP, and a choice between OpenAI-hosted and self-hosted execution environments.

That removes a meaningful amount of orchestration work, but it does not make the rest of the product disappear. agentbox would still need to decide what owns agent identity, durable memory, permissions, schedules, environment lifecycle, usage policy, and the user relationship.

This revision also adds a deliberately different product form: a Google Cloud Platform VM running Hermes Agent. Instead of consuming a managed harness, this approach assembles a persistent Linux computer and a relatively complete agent runtime into a self-operated agent appliance. Comparing the two makes the architectural choice more concrete.

The architectural shift ​

A useful working formula is:

text
agentbox
  = persistent agent identity
  + replaceable model intelligence
  + replaceable agent harness
  + durable memory and history
  + an on-demand private computer

GPT-6 Astra is a plausible intelligence layer for an experiment because it combines stronger reasoning with software engineering, computer use, browsing, tool orchestration, and long-running workflows. But the model is still not the durable agent.

The durable unit should survive model, harness, and computer replacement. It would own or reference at least:

  • an agent_id and owner;
  • profile and instructions;
  • permissions and policy;
  • credential references, never raw credentials in model context;
  • memory and conversation history;
  • tasks, schedules, and pending events;
  • model and harness configuration;
  • current execution-environment reference;
  • usage, limits, and billing state.

The practical principle is simple:

Identity, memory, permissions, history, and lifecycle belong to the product. Models, harnesses, and computers are replaceable implementations.

Agent and state layers ​

OpenAI Agents API ​

The Agents API is the most direct new baseline if GPT-6 Astra is the primary model. It divides the system into four concepts: a reusable agent configuration, an execution environment, a durable session, and the events and items produced while the agent works.

The significant addition is the managed Codex harness. OpenAI operates the model-and-tool loop, context compaction, recovery, steering, skills, MCP connections, and subagent orchestration. A session can continue across turns and can outlive the computer attached to it. The application can use an OpenAI-hosted sandbox or connect the same session to a self-hosted environment through the Codex executor.

This is close to the separation agentbox needs:

text
agentbox identity and policy
  -> OpenAI agent and durable session
  -> execution environment
       -> OpenAI-hosted sandbox
       -> agentbox-managed computer

It is not, however, a complete persistent-agent product. The application still owns the mapping between its user-facing agent_id, the OpenAI session, and the current environment. It also remains responsible for schedules and external events, product-level permissions, environment provisioning, backups, usage policy, billing, and recovery when a self-hosted executor disappears.

There are also material constraints to test rather than hide in an integration layer. The API is in public beta. Data residency is currently limited to the United States, zero data retention is not supported even with a self-hosted sandbox, and tracing cannot yet be exported through a customer-facing API. OpenAI-hosted sandboxes are working environments, not permanent computers: files can persist during the sandbox lifetime and selected outputs can survive expiry, but an inactive sandbox may be deleted after an hour.

The Agents API therefore reduces harness work while making the product boundary more important. agentbox should own a provider-neutral agent record and treat OpenAI agent and session identifiers as attached execution references, not as the canonical identity.

OpenAI Agents SDK ​

The Agents SDK remains useful when full control over the loop matters more than managed orchestration. It provides tools, handoffs, guardrails, tracing, and model context without requiring the Agents API's managed session model.

That makes it a valuable control case and an exit path, but no longer the obvious first implementation. Choosing it means agentbox must build or adopt the durable session, compaction, recovery, scheduling, and multi-agent machinery that the Agents API now supplies.

Letta ​

Letta starts closer to the persistent-agent side of the problem. Its agent state is stored and reconstructable, its memory blocks persist across interactions, and its current platform includes schedules, skills, long-lived message history, and a choice between managed and self-hosted deployment.

That makes Letta interesting in two roles:

  1. as the agent harness and memory system together;
  2. as a reference implementation for how agent identity and memory might be modeled.

The open question is ownership. If Letta becomes the canonical store for identity, memory, scheduling, and behavior, replacing it later becomes a data-model migration rather than a harness swap. A proof of concept should therefore test whether agentbox can keep a small canonical identity record while treating Letta state as an attached implementation.

Cloudflare Agents ​

Cloudflare Agents combines durable identity, embedded SQL state, real-time connections, scheduling, and recoverable workflows with adjacent browser and sandbox primitives. Its long-running-agent model is particularly relevant: an agent can wake, work, checkpoint, sleep, and later resume while its durable state and schedules remain intact.

This closely matches the idea of persistent identity paired with compute that scales down when idle. It could produce a compact implementation with few separately operated services.

The tradeoff is architectural concentration. Durable Objects, embedded storage, scheduling, workflows, browser execution, and deployment all reinforce one platform. That may be an excellent operational choice, but it should be made consciously because portability would require extracting several layers at once.

Meta Muse as a product reference ​

Muse combines a personal agent, a dedicated Linux VM, memory and files, browser and terminal work, schedules, connected services, and a user interface for goals, activity, and permissions. Meta says each VM is the system of record for its user's data. Muse is useful here as a benchmark for the complete experience rather than as an available harness or infrastructure provider.

Its security architecture also changes the comparison. The main agent works in an isolated runtime cell. Credential handling and built-in connectors run outside it, while a separate host-side Sentinel authorizes connector actions and network egress. Approval is delivered through the client independently of the agent conversation. A product claiming to give an agent a private computer should be evaluated on that authority boundary as well as on the computer itself.

These are Meta's published design claims, not a conclusion that Muse is immune to prompt injection or client compromise. Meta says its current service can access VM data when necessary to operate it; a Confidential VM intended to restrict that access was still forthcoming at this review date. A separate Mac client vulnerability disclosed in September shows why clients and approval channels belong in the threat model too.

Enterprise reference stacks ​

Microsoft Agent Framework ​

Microsoft Agent Framework is useful as a completeness benchmark. It covers agents, tools, MCP, memory and persistence, workflows, checkpoints, human approval, background execution, and hosting, while providing migration paths from AutoGen and Semantic Kernel.

For a small exploratory product, adopting the whole framework would likely bring more surface area than needed. Its value here is showing which concerns emerge once agent execution becomes an operated system rather than a demonstration.

Amazon Bedrock AgentCore ​

AgentCore divides the problem into modular services: Runtime, Memory, Gateway, Identity, Browser, Code Interpreter, and Observability. AWS explicitly separates its optional harness from the serverless runtime and supports third-party frameworks and models.

That separation is architecturally healthy and makes AgentCore another useful completeness benchmark. It may also be a future enterprise backend. For an early proof of concept, however, the operational and commercial commitment to AWS would need to earn its place against narrower suppliers.

Computers and hosting ​

The most consequential distinction in this category is not “container versus VM.” It is the persistence contract:

  • What survives idle suspension?
  • What survives a host failure or environment destruction?
  • Are memory and running processes preserved, or only files?
  • Can the environment be snapshotted, cloned, archived, and migrated?

Blaxel ​

Blaxel Sandboxes are VM-based environments designed for agent workloads. They advertise rapid resume, automatic scale-to-zero, filesystem and process restoration from standby, controlled networking, secrets, exposed ports, and API or MCP access.

There is an important qualification: standby state is not the same as guaranteed long-term storage. Blaxel's documentation recommends volumes for data that must survive sandbox destruction. A proof of concept should test both wake behavior and the recovery path when only the volume remains.

Blaxel is attractive when low wake latency and an agent-specific control plane matter more than infrastructure generality.

Daytona ​

Daytona offers container sandboxes, Linux and Windows VMs, and GPU environments with dedicated kernel, filesystem, network, compute, and disk boundaries. Its persistence semantics are unusually explicit:

  • stopped containers preserve their filesystem;
  • paused VMs preserve memory, running processes, and filesystem state;
  • snapshots, forks, volumes, and archives cover different durability and duplication needs.

This makes Daytona a strong candidate for testing both ordinary sandbox tasks and the stronger promise of a resumable personal computer. The proof of concept still needs to measure wake latency, regional availability, failure recovery, and cost under realistic idle patterns.

Fly Sprites ​

Fly describes Sprites as hardware-isolated Linux computers for agents. They use Firecracker isolation, retain a persistent filesystem, sleep when idle, wake quickly, and support checkpoint and restore.

Sprites are conceptually close to “give an AI a computer that stays alive.” They deserve evaluation if that becomes the product language, especially for workloads that expose services or accumulate a working environment over time. As with any newer platform surface, operational limits and migration paths need direct testing rather than inference from positioning.

E2B ​

E2B provides secure sandboxes with create, execute, pause, resume, and destroy lifecycle operations. Pausing preserves filesystem, memory, and process state, and snapshots support repeatable environments and checkpoints. Network connections do not survive a pause.

E2B remains a strong benchmark for coding, browser, and task-scoped execution. Its center of gravity is still the sandbox session rather than the permanent identity of a personal computer, so agentbox would need to provide that identity and continuity above it.

Modal Sandboxes are a good fit for high-volume, on-demand execution, parallel workers, batch jobs, and GPU workloads. Sandboxes are intentionally time-bounded; longer-lived work relies on filesystem or memory snapshots and later restoration.

That makes Modal more compelling as an elastic worker pool than as the default home for one persistent agent. It belongs in the architecture when a durable agent needs to fan out temporary work, not necessarily when it needs one private computer of record.

Google Cloud Platform VMs + Hermes Agent ​

A Google Compute Engine VM running Hermes Agent is not merely another sandbox supplier. It is a different product shape: a conventional private Linux computer combined with a relatively complete, self-hosted agent runtime.

text
Google Compute Engine VM
  -> Linux, persistent disk, network, IAM
  -> Hermes Agent
       -> model provider
       -> memory, skills, tools, schedules, channels

Google Cloud supplies familiar infrastructure primitives. A stopped VM can be started again with its persistent disks intact. Suspending a supported VM additionally preserves memory, but suspension has limits: it is not available for every machine type or accelerator configuration, and a VM cannot remain suspended indefinitely. Durable data should therefore live on persistent storage and have an independent backup path rather than depend on process memory.

Hermes supplies much of what would otherwise sit in the harness and application layers: a model-agnostic agent loop, persistent memory, skills, tools and MCP integrations, scheduled tasks, subagents, and gateways to messaging channels. It can run directly on a VPS and can also delegate terminal work to local, Docker, SSH, Daytona, Modal, and other backends.

Together, they form a useful thick agent appliance baseline. The strengths are legibility and control. Files and processes have ordinary Linux semantics; the model provider can be changed; debugging is direct; private services can stay on a stable network; and the product is not tied to a managed harness session.

The operational burden is equally concrete. agentbox would need to provision and harden the VM, patch the operating system and Hermes, manage secrets, monitor health, back up and restore state, implement wake and shutdown policy, and decide whether isolation is one VM per agent, per user, or per trust boundary. An always-running VM is simple but expensive at scale; an aggressively stopped VM is cheaper but needs an external control plane capable of waking it before Hermes can receive work.

Hermes memory and configuration should also be treated as an attached implementation until export and recovery have been tested. Installing a persistent agent on a persistent VM does not, by itself, separate the agent's identity from that machine.

Cloudflare Sandbox + Hermes Agent ​

The agentbox-cloud proof of concept uses a Cloudflare Worker and per-agent Durable Object to provision a Cloudflare Sandbox running Hermes. D1 indexes ownership, and an explicit backup action archives Hermes data and the workspace to R2. The Sandbox supplies an isolated Linux execution environment; the product supplies identity, lifecycle operations, and a dashboard.

This is a practical integrated form, but its present implementation does not yet meet the persistence and authority contract of a finished personal agent. Backups are manual and restore is not implemented. Model and Hermes API keys are passed into Sandbox environment variables for the runtime to use. There is no separate credential broker, outbound policy service, or user approval path comparable to Muse's Sentinel. Cloudflare's own Sandbox guidance distinguishes VM isolation from application authorization and notes that code inside a sandbox can read credentials supplied directly to its environment.

It is therefore a useful hands-on test of provisioning, service exposure, and state recovery. The next milestone is to prove restoration on a fresh environment and to define which actions and credentials must be controlled outside the agent runtime.

Product forms, not just suppliers ​

The supplier list becomes easier to reason about when the components are grouped into product forms:

Product formWhat the bundle ownsWhat agentbox still ownsMain tradeoff
Agents API + OpenAI-hosted sandboxManaged harness, session, model loop, temporary computerProduct identity, policy, external events, durable product stateFastest path, but the strongest OpenAI dependency and no permanent computer
Agents API + self-hosted environmentManaged harness and session; agentbox-selected computerProduct identity plus computer provisioning, recovery, and storageManaged orchestration with infrastructure choice, but split operational responsibility
GCP VM + Hermes AgentPrivate computer plus self-hosted harness, tools, memory, and channelsFleet control, identity boundary, security, backups, upgrades, billingMaximum legibility and control, but the largest operating surface
Cloudflare Sandbox + Hermes AgentIsolated computer, Hermes runtime, and an application-operated control planeState restoration, credential boundary, action policy, approvals, fleet operationsConcrete integrated PoC, with substantial product-level security and recovery work remaining
Letta + replaceable sandboxPersistent-agent and memory layer plus selected executionProduct identity boundary and provider lifecycleStrong state model, with migration risk if Letta becomes canonical
Cloudflare AgentsDurable runtime, state, schedules, connections, and adjacent executionProduct semantics and export strategyCompact operating model, with several layers concentrated on one platform
Meta MuseComplete personal agent, dedicated VM, connected services, permissions, and user experienceNo component is offered here as an agentbox backendConsumer-product benchmark; full agent portability is not publicly demonstrated

These are not interchangeable hosting plans. They place product responsibility at different layers. A useful proof of concept should compare the amount of product behavior each form supplies, the amount agentbox must operate, and how much state can be moved out later.

Working shortlist ​

LayerFirst candidatesAlternatives or benchmarks
Primary modelGPT-6 AstraClaude, Gemini, future models
Managed agent harnessOpenAI Agents APILetta
Self-hosted agent harnessHermes AgentOpenAI Agents SDK
Persistent memory and stateagentbox-owned schema, LettaCloudflare embedded state
Durable runtimeCloudflare Agentsself-operated control plane
Agent computerGCP Compute Engine, Blaxel, DaytonaFly Sprites
Integrated agent appliance PoCCloudflare Sandbox + Hermes AgentGCP VM + Hermes Agent
Task-scoped executionE2BModal
Enterprise referenceAWS AgentCoreMicrosoft Agent Framework
Consumer-product referenceMeta MuseNot an infrastructure supplier

This table is not a ranking. It identifies a small set worth testing and a wider set useful for comparison.

Five proof-of-concept architectures ​

A. Open and replaceable ​

text
agentbox identity, policy, lifecycle, and billing
  -> OpenAI Agents API + GPT-6 Astra
  -> self-hosted environment adapter
       -> GCP Compute Engine
       -> Blaxel or Daytona

agentbox keeps the canonical agent record, maps it to an OpenAI agent and session, and exposes a narrow environment interface: create, connect, wake, execute, checkpoint, stop, destroy, attach storage, and inspect usage.

This is a clean test of the central thesis while avoiding the need to build the entire harness. Can the same agent identity and session move between two computers without changing what the user believes the agent is?

The cost is a divided control plane. OpenAI owns the harness and session; agentbox owns the computer lifecycle and must reconcile both systems. Scheduling, external event delivery, durable product state, and migrations also remain product responsibilities.

B. Persistent agent and memory ​

text
agentbox account and product layer
  -> Letta persistent agent and memory
  -> GPT-6 Astra or another model
  -> replaceable computer provider

This option tests whether a mature stateful-agent layer removes enough complexity to justify the dependency. A useful metaphor is identity and memory as the durable mind, with a replaceable model and computer as reasoning engine and body.

The critical experiment is exportability: whether memory, history, schedules, and behavioral state can be reconstructed outside Letta without semantic loss.

C. Cloudflare-native ​

text
Cloudflare Agent identity + SQLite state
  -> schedules and durable workflows
  -> browser, sandbox, or container execution
  -> GPT-6 Astra

This is likely the smallest operated system. Durable identity, state, connections, schedules, and recoverable work live in one runtime that naturally sleeps between events.

It should be evaluated as a coherent platform choice, not as a set of easily replaceable components. The proof of concept should include an explicit export and recovery exercise to quantify lock-in rather than discussing it abstractly.

D. Full agent appliance ​

text
agentbox control plane
  -> one GCP VM per agent, account, or trust boundary
       -> Hermes Agent
            -> selected model provider
            -> memory, skills, schedules, tools, and channels

This option tests the most concrete expression of “give an AI a private computer.” The VM is the durable home and Hermes is the resident agent runtime. It is easy to inspect and does not require a separate managed harness, making it a useful control against more abstract platforms.

The experiment should not stop when the agent answers a message. It must include automated provisioning, secure bootstrap, stop and start behavior, health checks, upgrades, backup and restore, model switching, and reconstruction of the same agent on a fresh VM. That last exercise determines whether the agent is genuinely portable or merely a carefully maintained server.

E. Cloudflare Sandbox with Hermes ​

text
agentbox-cloud Worker + account index
  -> per-agent Durable Object lifecycle
  -> Cloudflare Sandbox running Hermes
  -> R2 backup of agent data and workspace

This is the active code-level proof of concept; the real Cloudflare runtime has not yet passed end-to-end integration testing. It tests whether a managed container substrate can support a ready-to-use agent computer with a small control plane. It should be judged against its actual implementation: a backup button is not yet recovery, Worker secrets injected into the Sandbox are visible to code in that environment, and account authentication does not govern the agent's subsequent network actions.

The next test should restore the same agent on fresh compute, then add a narrowly scoped action and credential boundary outside Hermes. Muse supplies a demanding reference for the second test without prescribing its exact implementation.

Evaluation criteria ​

The first round should test behavior, not feature-list parity.

Identity and agent state ​

  • Can an agent keep the same identity while its model changes?
  • Can it move to a new harness or computer without losing memory or pending work?
  • Are history, memory, schedules, and task state exportable?
  • Can a managed session be reconstructed outside its original API?
  • Can context be compacted without silently changing durable state?
  • Can tools, MCP servers, and credentials be granted and revoked independently?
  • Are human approval, cancellation, retry, and audit trails first-class?

Authority and safety ​

  • Can agent-generated code reach the network without passing a policy decision outside its runtime?
  • Can the agent read a real credential, or does a broker supply it only for an approved action?
  • Are approvals bound to the specific action, destination, and duration, and delivered outside the agent conversation?
  • Can the user inspect and revoke permissions independently of the agent's own memory or instructions?
  • Are browser sessions, local clients, and exposed services included in the same threat model?

Computer lifecycle ​

  • What survives stop, suspend, pause, provider failure, and destruction?
  • How long do cold start and warm resume take at realistic percentiles?
  • Are filesystem, process memory, network identity, and running services preserved?
  • Can environments be snapshotted, cloned, archived, and restored?
  • Are outbound network policy, static egress, secrets, browser access, ports, and GPU workloads supported?
  • Does durable storage have an independent backup and recovery path?
  • Can the same environment be attached to another harness without rebuilding it?

Operations and commercial fit ​

  • What is the true cost of active, idle, stored, and resumed environments?
  • Which quotas, time limits, regions, and concurrency caps apply?
  • Are logs, traces, resource usage, and failure reasons observable?
  • What availability commitments and support channels exist?
  • Can the architecture be migrated or self-hosted if the supplier changes direction?
  • Who patches the runtime and operating system, and how are failed upgrades rolled back?

Pricing should be measured with a workload model, not copied into this note. Price sheets, limits, and product boundaries change faster than the underlying architectural questions.

Product implication ​

The old aClaw proposition was close to:

Rent an OpenClaw or Hermes agent.

That language binds the product to an implementation. A more durable proposition would be one of these:

Deploy a persistent AI agent.

or:

Give an AI a private computer that can wake up and continue its work.

The second is more concrete, but the computer alone is still insufficient. The product value is the continuity across computers:

text
identity
+ memory
+ permissions
+ history
+ lifecycle
+ execution

That is the layer agentbox should understand before it decides whether there is another product to build.

The two newly compared forms expose the choice clearly. OpenAI Agents API lets agentbox start with a thin product layer above a managed harness. GCP plus Hermes starts with a thick, inspectable agent appliance under an agentbox control plane. Neither answer is universally better. The former optimizes for speed and managed orchestration; the latter optimizes for ownership, portability of model access, and conventional systems control.

Muse shows what the assembled consumer experience can look like. Its dedicated computer supports work over time, while its separate authority governs what the agent may do with access to the user's wider digital life. The agentbox-cloud PoC can test the computer and control-plane parts now; the permission and recovery parts remain explicit work.

The strategically useful boundary sits above both:

text
agentbox identity, policy, lifecycle, and user relationship
  -> managed agent session
     or
  -> self-hosted agent appliance

Working conclusion ​

The next agentbox runtime, if there is one, should be modular from the beginning.

Models should be replaceable. Agent harnesses should be replaceable. Computers and hosting suppliers should be replaceable. Durable identity, permissions, history, lifecycle, and the user relationship should remain under the product's control.

The supplier set worth hands-on evaluation remains deliberately small:

  1. OpenAI Agents API for a managed GPT-6 Astra harness and durable session;
  2. GCP Compute Engine plus Hermes Agent for a self-hosted agent appliance;
  3. Letta for persistent agent state and memory;
  4. Blaxel and Daytona for private, resumable computers;
  5. Cloudflare Agents for a tightly integrated durable-runtime alternative.

The OpenAI Agents SDK remains the lower-level control case when agentbox needs to own the loop. AWS AgentCore and Microsoft Agent Framework provide enterprise reference points. E2B and Modal remain useful for task-scoped execution. Fly Sprites is worth adding when the experiment specifically tests the promise of a persistent personal Linux computer.

Cloudflare Sandbox plus Hermes is now an active integrated PoC, distinct from Cloudflare Agents. Muse is a consumer-product reference rather than an infrastructure candidate. Both sharpen the acceptance criteria: continuity must survive environment replacement, and permission must be enforced outside the agent's own execution cell.

The central thesis is unchanged by any supplier choice:

An agent is not a model, a process, or a container. It is the persistent identity and state that can continue across all three.

Sources and freshness ​

The earlier supplier survey was checked on September 14, 2026. Muse and Cloudflare Sandbox claims were checked against official documentation on September 24, 2026. All capabilities should be rechecked before an implementation or purchasing decision.

An ongoing exploration of infrastructure for AI agents. Released under the Apache 2.0 License.