Agent Runtime

Bring your harness.
Keep your infrastructure.

Thanks for joining!
ii-agent-runtime · developer preview

$ ii-agent-runtime deploy aws codex-production

setup rev 13 · sha256:7c9e41b2 · codex 0.42.0 · gpt-5-codex · repo@5034eb3

access 5 credential refs checked live · 0 escalation paths · prod unreachable

launch alice@acme.com authorized by your identity provider

infra runs in aws

deployed codex-production rev 13

$ ii-agent-runtime run codex-production --task sympy__sympy-20590

run-043 42.8s · 182.4k in · 6.9k out · $0.13 · trace, I/O and outputs kept

evaluated SWE-bench Verified with your evaluator: resolved · rerun from rev 13

Prevent agent privilege escalation.

Configure GitHub, AWS, Google Cloud, Azure, or custom credentials and see everything your harness can and can't do.

permission check · codex-production rev 12

live check · analyzed as ii-analyzer (read-only) · target credentials not elevated

refs github-app/ii-runtime-prod · arn:aws:iam::123456789012:role/ii-agent

ii-agent@acme-prod.iam.gserviceaccount.com · userAssignedIdentities/ii-agent

vault://secret/ii/billing-api

can 47 actions contents:write pull_requests:write s3:GetObject ...

cannot 12,406 actions secrets:read iam:CreateAccessKey roleAssignments/write ...

BLOCK escalation aws iam:PassRole + ec2:RunInstances -> role/deploy-admin

BLOCK indirect github contents:write can push a branch that runs deploy.yml with PROD_TOKEN

BLOCK regression gcp storage.objects.delete reachable since rev 11

WARN excessive aws s3:DeleteBucket granted, never needed by this setup

WARN redundant azure Storage Blob Data Reader already covered by Blob Data Contributor

PASS assertion must not reach Key Vault kv-prod

PASS custom test billing-api refunds endpoint unreachable

deployment gated: 3 blocking findings

Make every run reproducible.

Every run stays tied to the exact setup that produced it, so you can see what changed and recreate it.

deploy.json

{

"setup": "codex-production",

"revision": 12,

"harness": "@openai/codex@0.42.0",

"workspace": {

"repository": "intelligent-iterations/ii-agent-runtime",

"commit": "5034eb3dfe3829edb3282d327592735c7b1193b2"

},

"runtime": {

"image": "ghcr.io/intelligent-iterations/agent-base",

"digest": "sha256:1d4028a6ede6bebe8f49c3ab039b22b8a2e785f854e25acaa7ca9073889cc758",

"cpu": "4",

"memory": "8Gi",

"region": "us-east-1"

}

}

Turn runs into evidence.

Capture the full execution so you can trace what happened, measure what it cost, and evaluate the result with SWE-bench, Harbor, or your own evaluator.

run-042 · execution record

trace 19518d6f1a9d2edf531e4c8df9f69fd8 · run-042 · sympy__sympy-20590

invoke_agent codex 42.8s in 182.4k (151.0k cached) out 6.9k $0.13

chat gpt-5-codex 9.6s

execute_tool shell 1.1s grep -rn "__slots__" sympy/core

execute_tool apply_patch 0.4s sympy/core/_print_helpers.py +5

execute_tool shell 18.2s pytest sympy/core/tests/test_basic.py exit 0

exported run-042.otlp.json · predictions.json

$ swebench eval verified -p predictions.json --run-id run-042

Instances resolved: 1

How teams run coding agents in production

Keep your models, identities, credentials, compute, and telemetry storage. We’re building the reusable infrastructure around them. Use the pieces you need; extend the rest.

Thanks for joining!