$ ii-agent-runtime deploy aws codex-production
setup rev 13 · sha256:7c9e41b2 · codex 0.42.0 · gpt-5-codex · repo@5034eb3
access 5 credential refs checked live · 0 escalation paths · prod unreachable
launch alice@acme.com authorized by your identity provider
infra runs in aws
✓ deployed codex-production rev 13
$ ii-agent-runtime run codex-production --task sympy__sympy-20590
✓ run-043 42.8s · 182.4k in · 6.9k out · $0.13 · trace, I/O and outputs kept
✓ evaluated SWE-bench Verified with your evaluator: resolved · rerun from rev 13
01 / Secure
Prevent agent privilege escalation.
Configure GitHub, AWS, Google Cloud, Azure, or custom credentials and see everything your harness can and can't do.
live check · analyzed as ii-analyzer (read-only) · target credentials not elevated
refs github-app/ii-runtime-prod · arn:aws:iam::123456789012:role/ii-agent
ii-agent@acme-prod.iam.gserviceaccount.com · userAssignedIdentities/ii-agent
vault://secret/ii/billing-api
can 47 actions contents:write pull_requests:write s3:GetObject ...
cannot 12,406 actions secrets:read iam:CreateAccessKey roleAssignments/write ...
BLOCK escalation aws iam:PassRole + ec2:RunInstances -> role/deploy-admin
BLOCK indirect github contents:write can push a branch that runs deploy.yml with PROD_TOKEN
BLOCK regression gcp storage.objects.delete reachable since rev 11
WARN excessive aws s3:DeleteBucket granted, never needed by this setup
WARN redundant azure Storage Blob Data Reader already covered by Blob Data Contributor
PASS assertion must not reach Key Vault kv-prod
PASS custom test billing-api refunds endpoint unreachable
deployment gated: 3 blocking findings
02 / Deploy
Make every run reproducible.
Every run stays tied to the exact setup that produced it, so you can see what changed and recreate it.
{
"setup": "codex-production",
"revision": 12,
"harness": "@openai/codex@0.42.0",
"workspace": {
"repository": "intelligent-iterations/ii-agent-runtime",
"commit": "5034eb3dfe3829edb3282d327592735c7b1193b2"
},
"runtime": {
"image": "ghcr.io/intelligent-iterations/agent-base",
"digest": "sha256:1d4028a6ede6bebe8f49c3ab039b22b8a2e785f854e25acaa7ca9073889cc758",
"cpu": "4",
"memory": "8Gi",
"region": "us-east-1"
}
}
03 / Measure
Turn runs into evidence.
Capture the full execution so you can trace what happened, measure what it cost, and evaluate the result with SWE-bench, Harbor, or your own evaluator.
trace 19518d6f1a9d2edf531e4c8df9f69fd8 · run-042 · sympy__sympy-20590
invoke_agent codex 42.8s in 182.4k (151.0k cached) out 6.9k $0.13
chat gpt-5-codex 9.6s
execute_tool shell 1.1s grep -rn "__slots__" sympy/core
execute_tool apply_patch 0.4s sympy/core/_print_helpers.py +5
execute_tool shell 18.2s pytest sympy/core/tests/test_basic.py exit 0
exported run-042.otlp.json · predictions.json
$ swebench eval verified -p predictions.json --run-id run-042
Instances resolved: 1
Make it part of your project
How teams run coding agents in production
Keep your models, identities, credentials, compute, and telemetry storage. We’re building the reusable infrastructure around them. Use the pieces you need; extend the rest.
Thanks for joining!Couldn't save that. Please try again.

