,

AgentCore Runtime Instances: Engineering Guide

A robotic AI agent figure standing beside a bank of physical GPU server racks

AgentCore Runtime Instances, generally available as of August 6, 2026, are Amazon’s answer to a problem every team running production agents eventually hits: some workloads just don’t fit in an 8-hour box. Long research loops, multi-day document processing, anything that needs a GPU sitting warm and ready — the old AgentCore runtime handled none of that gracefully. Runtime instances fix it by putting your agent on real EC2 capacity that AgentCore itself provisions, patches, and scales, without you standing up an ASG, writing a launch template, or getting paged when an instance fails a health check. Let’s get into the actual mechanics.

What AgentCore Runtime Instances actually are

Up to now, Bedrock AgentCore ran agents on managed microVMs — fast to start, fully isolated, but capped at 8 hours of runtime per session. Fine for a chat-style agent or a bounded task. Useless for an agent that’s chewing through a multi-day data pipeline or running inference against a fine-tuned model that needs a GPU parked under it the whole time. Runtime instances add a second execution mode: your agent runs on dedicated EC2 instances that AgentCore manages on your behalf. You’re not SSHing into boxes or babysitting an Auto Scaling group — you’re declaring what kind of compute your agent needs and letting AgentCore handle provisioning, patching, scaling, and lifecycle from there. Think of it as the managed middle ground between “fully serverless microVM” and “roll your own EC2 fleet.”

Capacity providers: the contract between your agent and its compute

The core new primitive is the capacity provider. A capacity provider is a declaration of which EC2 instance types or families your agents are allowed to run on — GPU-accelerated families like the G6 or P5 line, memory-optimized families like R7g for context-heavy workloads, or compute-optimized families like C7g when you’re doing more orchestration than inference. You create the capacity provider once, then attach one or more agents to it. AgentCore reads that contract and handles the rest: spinning up instances that match, scaling within your bounds, and rotating capacity as instances age out or need patching.

aws bedrock-agentcore create-capacity-provider \
  --capacity-provider-name gpu-inference-pool \
  --instance-families '["g6.2xlarge","g6.4xlarge","g6.8xlarge"]' \
  --min-instances 1 \
  --max-instances 8 \
  --scaling-policy TargetUtilizationPercent=70

aws bedrock-agentcore update-agent-runtime-config \
  --agent-runtime-id doc-analysis-agent \
  --runtime-mode INSTANCE \
  --capacity-provider-name gpu-inference-pool \
  --session-ttl-days 14

That’s the whole handshake: define the capacity provider, point an agent’s runtime config at it, set a session TTL. No launch templates, no target tracking policies to hand-tune, no AMI pipeline to maintain — AgentCore owns the EC2 layer underneath the abstraction. Full parameter reference lives in the AgentCore documentation.

Picking the right instance family

This is where runtime instances earn their keep over the microVM path — you get to choose hardware that actually matches the workload instead of accepting a generic sandbox. GPU-accelerated families make sense for agents doing on-instance inference, embedding generation at volume, or any task where a model call needs local acceleration rather than a round trip to a separate endpoint. Memory-optimized families are the right call for agents holding large context windows, big in-memory indexes, or multi-agent state that needs to stay resident across a long session. Compute-optimized families fit orchestration-heavy agents — the ones spending most of their cycles on tool calls, branching logic, and coordinating other agents rather than running models directly. You can mix families within a single capacity provider’s allowed list, and AgentCore will place workloads accordingly, so you’re not forced into one instance type for every agent hanging off that provider. Full family specs are in the EC2 instance types documentation.

Session persistence: 14 days instead of 8 hours

The microVM runtime’s 8-hour ceiling was a hard wall for anything resembling a long-running workflow — a multi-day research agent, a batch job chewing through a large document set, a monitoring agent that needs to keep state across a weekend. Runtime instances support session persistence up to 14 days, which changes what’s actually viable to build. State survives across that window, so an agent can pick up exactly where it left off instead of you engineering a checkpoint-and-resume layer just to work around a timeout. Multi-agent collaboration on a shared host is also supported here — several agents can run against the same instance capacity, useful when you’ve got a coordinator agent and several worker agents that genuinely benefit from being co-located rather than talking over the network for every handoff.

A humanoid AI assistant icon loading a shipping container box onto a server rack

Containerized deployment for teams that want to ship independently

Runtime instances also support containerized agent deployment, which matters more than it sounds like on paper. If your team is already building agents as containers — packaging dependencies, pinning model client versions, running the same image through CI that you’ll run in production — you don’t have to change that workflow to adopt runtime instances. You ship a container, AgentCore places it on matching capacity from your capacity provider, and your team keeps shipping on its own cadence instead of waiting on a shared deployment pipeline. That’s a meaningful difference from the microVM runtime’s more opinionated packaging model, and it’s the detail that will matter most if you’ve got several teams running agents that don’t want to coordinate releases with each other.

Cost control: stop and restart sessions

Long-running doesn’t have to mean always-billing. Runtime instances support session stop and restart, so an agent sitting idle between work — waiting on an external event, paused overnight, whatever the case — doesn’t have to burn compute the whole time. You stop the session, state persists, you restart it later and pick up where you left off. That’s the lever that makes 14-day persistence economically sane instead of an implicit invitation to leave GPU instances running for two weeks straight.

Runtime instances vs. the microVM runtime: picking the right lane

Nobody should be migrating everything to runtime instances just because they’re new. The microVM runtime is still the right default for short, bursty, stateless-ish agent invocations — it starts fast, isolates cleanly, and you pay for exactly the seconds you use, capped at 8 hours. Reach for runtime instances when a workload has any of these traits: it needs to run longer than 8 hours, it needs a GPU or a large memory footprint parked under it, it needs multiple agents genuinely co-located rather than just talking over an API, or your team wants container-level deployment control instead of the managed microVM packaging. In practice, most organizations will run both side by side — microVM for the high-volume, short-lived agent traffic, runtime instances for the handful of workloads that were previously impossible to run inside AgentCore at all. Background on the original runtime model is still worth a read in the AWS Machine Learning Blog and the AWS News Blog coverage of the AgentCore launch.

What this actually unblocks

The honest read here: this isn’t AWS reinventing agent compute, it’s AWS admitting that a single 8-hour execution model was never going to cover every agent workload, and building the managed EC2 path instead of leaving you to hand-roll one. If you’ve been quietly running a side fleet of EC2 instances just to handle the agents that didn’t fit AgentCore’s original constraints, runtime instances are the signal to fold that fleet back into a single managed surface.

At Enkompass (enkompass.net), this is our home turf. We’ve shipped production Bedrock and AgentCore workloads that actually move the needle, not just demo well. If you’d rather have seasoned cloud people handle it instead of learning the hard way at 2 AM, let’s talk. Reach out through our contact page or call us directly at +1-828-436-5667 — we’ll help you turn this from a someday-project into a done-deal.