feat(migration): phase 4 — sandbox capability + healthcheck + session manager

Three modules under strix/sandbox/ that bring the per-scan container
plumbing in line with the SDK's capability model:

- healthcheck.py: wait_for_http_ready (FastAPI tool server /health)
  and wait_for_tcp_ready (Caido proxy port — no /health endpoint).
  Connect/timeout errors continue polling; the timeout error message
  carries the last failure class so a stuck scan tells you whether the
  port refused, hung, or returned a non-2xx.

- caido_capability.py: CaidoCapability subclasses agents.sandbox.
  capabilities.Capability and wires three concerns:
  1. process_manifest injects http_proxy / https_proxy / ALL_PROXY
     env vars pointing at the in-container Caido listener.
  2. tools() returns the seven Caido SDK function tools from Phase 2.5
     so the SDK runtime auto-merges them with each agent's tool list.
  3. bind() schedules an asyncio.gather of both healthcheck probes;
     StrixOrchestrationHooks.on_agent_start awaits the resulting
     task before the first LLM call.
  Pydantic v2 PrivateAttr is used for the underscore-prefixed runtime
  fields (Pydantic forbids underscore-prefixed model fields).

- session_manager.py: per-scan_id cache. create_or_reuse builds the
  StrixDockerSandboxClient with docker.from_env() (the SDK's docker
  client now requires an explicit DockerSDKClient instance at init),
  constructs the Manifest via Environment(value=...) (a flat dict is
  silently dropped by Pydantic), resolves the host-side mapped ports
  via session._resolve_exposed_port, configures the capability with
  those ports *before* binding, and returns a bundle dict the
  per-agent context reads to populate tool_server_host_port /
  caido_host_port / bearer. cleanup is best-effort: a Docker daemon
  error during delete is logged and swallowed so a stranded
  container doesn't block the next scan.

Tests: 21 new tests in tests/sandbox/ — healthcheck happy path /
polling-through-failures / timeout for both HTTP and TCP probes (the
TCP test uses a real local listener, no mocks); CaidoCapability env
injection / tool list / bind scheduling / configure_host_ports;
session_manager full create flow, cache reuse, custom timeout, cleanup
including the Docker-daemon-failure swallow path.

mypy override added for docker.* (no upstream stubs); per-file ruff
TC002 ignore added for caido_capability.py — agents.tool.Tool is used
at runtime for the cached _CAIDO_TOOLS tuple.

Refs: PLAYBOOK.md §3.1-3.3, AUDIT.md §2.5 (C5).
This commit is contained in:
0xallam
2026-04-25 00:49:26 -07:00
parent b5578007c4
commit 775a78487d
8 changed files with 1074 additions and 0 deletions
+212
View File
@@ -0,0 +1,212 @@
"""CaidoCapability — sandbox capability for the Caido HTTP/HTTPS proxy.
Three concerns wired into the SDK's capability lifecycle:
1. **Manifest mutation** (``process_manifest``): inject ``http_proxy`` /
``https_proxy`` / ``ALL_PROXY`` env vars pointing at the in-container
Caido listener. Any tool that ultimately shells out (curl, requests,
etc.) now flows through the proxy automatically.
2. **Tool exposure** (``tools``): the seven Caido SDK function-tool
wrappers from Phase 2.5 are returned here. The SDK runtime collects
tools from every capability and merges them with the agent's
``tools=[...]`` declaration, so individual agents don't have to
redeclare them.
3. **Healthcheck task** (``bind``): when a session binds, we kick off
:func:`wait_for_http_ready` against the FastAPI tool server's
``/health`` endpoint and :func:`wait_for_tcp_ready` against the
Caido proxy port. The aggregated task handle is stored on
``self._healthcheck_task``, which the
:class:`StrixOrchestrationHooks.on_agent_start` hook awaits before
the first LLM call so the agent never hits a connection-refused
on its very first tool invocation.
References:
- PLAYBOOK.md §3.2
- AUDIT.md §2.5 (C5 — healthcheck wired to RunHooks)
"""
from __future__ import annotations
import asyncio
import logging
from typing import TYPE_CHECKING, ClassVar, Literal
from agents.sandbox.capabilities.capability import Capability
from agents.tool import Tool
from pydantic import PrivateAttr
from strix.sandbox.healthcheck import wait_for_http_ready, wait_for_tcp_ready
from strix.tools.proxy.proxy_sdk_tools import (
list_requests,
list_sitemap,
repeat_request,
scope_rules,
send_request,
view_request,
view_sitemap_entry,
)
if TYPE_CHECKING:
from agents.sandbox.manifest import Manifest
from agents.sandbox.session.base_sandbox_session import BaseSandboxSession
logger = logging.getLogger(__name__)
# Container-internal Caido listener. The in-container Caido sidecar binds
# on this port; the host gets a randomly mapped port we resolve at
# session create time and pass into the per-agent context as
# ``caido_host_port`` for the proxy SDK tools' dispatcher.
_CAIDO_INTERNAL_PORT = 48080
# Container-internal FastAPI tool server. Same shape as Caido — host
# port is resolved at session create.
_TOOL_SERVER_INTERNAL_PORT = 48081
# Probe URLs used inside ``bind``. ``host=127.0.0.1`` because the host
# port mapping is loopback-only (legacy and SDK both bind to 127.0.0.1).
_PROBE_HOST = "127.0.0.1"
# Cached tool list — building Tool instances has side effects via the
# function_tool decorator and we don't want re-instantiation each time
# the SDK calls ``tools()``.
_CAIDO_TOOLS: tuple[Tool, ...] = (
list_requests,
view_request,
send_request,
repeat_request,
scope_rules,
list_sitemap,
view_sitemap_entry,
)
class CaidoCapability(Capability):
"""Caido HTTP/HTTPS forward proxy + 7 GraphQL function tools.
Lifetime: one instance per scan. The SDK clones capabilities
per-run (see ``Capability.clone``); we accept that — each cloned
instance opens its own healthcheck task on ``bind``, which is
cheap and idempotent.
"""
type: Literal["caido"] = "caido"
# Pydantic ``PrivateAttr`` for runtime-only state. Pydantic forbids
# underscore-prefixed *fields*, but private attributes are first-class
# and cleanly excluded from model dumps and serialization.
_healthcheck_task: asyncio.Task[None] | None = PrivateAttr(default=None)
# The two ports the host needs to reach. Populated by the session
# manager *after* the SDK creates the container and we've resolved
# the random host-side mappings via ``session._resolve_exposed_port``.
_tool_server_host_port: int | None = PrivateAttr(default=None)
_caido_host_port: int | None = PrivateAttr(default=None)
# Per-capability healthcheck timeout. Long enough to cover image
# pulls on a cold cache plus tool-server boot, short enough that a
# mis-configured image fails the run inside a few minutes.
_HEALTHCHECK_TIMEOUT: ClassVar[float] = 60.0
def process_manifest(self, manifest: Manifest) -> Manifest:
"""Inject proxy env vars into the manifest's environment.
Mutates in place; returns the same manifest. Mirrors the SDK's
Capability protocol where ``process_manifest`` is the single
synchronous hook for changing what the container sees.
"""
env = dict(manifest.environment.value or {})
env.update(
{
"http_proxy": f"http://127.0.0.1:{_CAIDO_INTERNAL_PORT}",
"https_proxy": f"http://127.0.0.1:{_CAIDO_INTERNAL_PORT}",
"ALL_PROXY": f"http://127.0.0.1:{_CAIDO_INTERNAL_PORT}",
},
)
manifest.environment.value = env
return manifest
def tools(self) -> list[Tool]:
"""Return the seven Caido function tools.
The SDK runtime calls this at agent-build time and merges the
result with the agent's own tool list. Returning a fresh list
each call (rather than yielding the cached tuple directly) is
SDK convention.
"""
return list(_CAIDO_TOOLS)
async def instructions(self, manifest: Manifest) -> str | None: # noqa: ARG002
"""System-prompt fragment appended for every Caido-equipped agent."""
return (
"<caido_proxy>\n"
"All HTTP/HTTPS traffic in this sandbox is automatically captured "
f"by Caido (in-container at 127.0.0.1:{_CAIDO_INTERNAL_PORT}; "
"host_proxy / https_proxy env vars are pre-set).\n"
"Tools: list_requests, view_request, send_request, repeat_request, "
"scope_rules, list_sitemap, view_sitemap_entry.\n"
"HTTPQL filter examples: "
"'request.method == \"POST\"', "
"'response.status >= 400', "
"'request.host == \"target.com\"'.\n"
"</caido_proxy>"
)
def configure_host_ports(
self,
*,
tool_server_host_port: int,
caido_host_port: int,
) -> None:
"""Record the resolved host-side ports.
Called by the session manager after ``client.create(...)``
returns, before binding the session. The healthcheck task
reads these to know which mapped ports to probe.
"""
self._tool_server_host_port = tool_server_host_port
self._caido_host_port = caido_host_port
def bind(self, session: BaseSandboxSession) -> None:
"""Schedule a healthcheck task on session bind.
Stores the task handle so :class:`StrixOrchestrationHooks` can
await it on the first agent start. We never raise from here —
the healthcheck failure surfaces inside on_agent_start, which
is the right place to fail the run because by then we have a
live RunContextWrapper to log against.
"""
super().bind(session)
if self._tool_server_host_port is None or self._caido_host_port is None:
logger.warning(
"CaidoCapability.bind called before configure_host_ports; "
"skipping healthcheck task scheduling.",
)
return
self._healthcheck_task = asyncio.create_task(
self._run_healthcheck(),
name=f"caido-healthcheck-{self._tool_server_host_port}",
)
async def _run_healthcheck(self) -> None:
"""Probe both ports concurrently; raise on first failure."""
# Mypy sees these as Optional, but ``bind`` checks both before
# creating the task.
assert self._tool_server_host_port is not None
assert self._caido_host_port is not None
await asyncio.gather(
wait_for_http_ready(
f"http://{_PROBE_HOST}:{self._tool_server_host_port}/health",
timeout=self._HEALTHCHECK_TIMEOUT,
),
wait_for_tcp_ready(
_PROBE_HOST,
self._caido_host_port,
timeout=self._HEALTHCHECK_TIMEOUT,
),
)
+121
View File
@@ -0,0 +1,121 @@
"""Sandbox port readiness probes used during session bring-up.
The in-container tool server (FastAPI) takes a few seconds to start
listening after the Docker container is created, and Caido's HTTPS
proxy takes a similar window. The session manager waits for both
before returning a session bundle so that the first tool call from
an agent doesn't hit a connection refused.
Two helpers are exposed:
- :func:`wait_for_http_ready` for the FastAPI tool server, whose
``/health`` endpoint returns ``{"status": "healthy"}`` once the
process is up. We don't require the JSON shape exactly — any 2xx
is treated as ready, mirroring the legacy ``_wait_for_tool_server``
but more lenient (the legacy version checked the JSON body too,
which made test images without that handler fail spuriously).
- :func:`wait_for_tcp_ready` for Caido, which serves an HTTP forward
proxy on its port and does *not* expose ``/health``. A TCP connect
is the most we can probe without sending real proxy traffic.
References:
- PLAYBOOK.md §3.1
- HARNESS_WIKI.md §6.4 (legacy ``_wait_for_tool_server`` pattern)
"""
from __future__ import annotations
import asyncio
import contextlib
import logging
import httpx
logger = logging.getLogger(__name__)
class SandboxNotReadyError(Exception):
"""Raised when a sandbox port doesn't accept connections in time."""
# Default per-attempt HTTP timeout. The legacy harness used 5s; we
# match it so a slow first request (image still warming up) doesn't
# misfire as a hard failure on a single attempt.
_DEFAULT_HTTP_PROBE_TIMEOUT = 5.0
# Default polling cadence between attempts. Balanced for CI-style
# fast bring-up (sub-second) without burning CPU when the port is
# legitimately taking a few seconds.
_DEFAULT_POLL_INTERVAL = 0.5
async def wait_for_http_ready(
url: str,
*,
timeout: float = 30.0,
poll_interval: float = _DEFAULT_POLL_INTERVAL,
probe_timeout: float = _DEFAULT_HTTP_PROBE_TIMEOUT,
) -> None:
"""Poll ``url`` until any 2xx response, or raise after ``timeout``.
Network errors (ConnectError / TimeoutException / RequestError)
are treated as "not ready yet" — the loop continues. Any other
exception class will surface immediately so a programmer error
(bad URL, etc.) doesn't get silently retried for 30 seconds.
"""
deadline = asyncio.get_event_loop().time() + timeout
last_error: str | None = None
async with httpx.AsyncClient(timeout=probe_timeout, trust_env=False) as client:
while asyncio.get_event_loop().time() < deadline:
try:
response = await client.get(url)
if 200 <= response.status_code < 300:
return
last_error = f"HTTP {response.status_code}"
except (httpx.ConnectError, httpx.TimeoutException, httpx.RequestError) as e:
last_error = type(e).__name__
await asyncio.sleep(poll_interval)
raise SandboxNotReadyError(
f"HTTP probe of {url} did not return 2xx within {timeout}s (last error: {last_error})",
)
async def wait_for_tcp_ready(
host: str,
port: int,
*,
timeout: float = 30.0,
poll_interval: float = _DEFAULT_POLL_INTERVAL,
) -> None:
"""Poll ``host:port`` until a TCP connect succeeds, or raise after ``timeout``.
Used for ports that don't expose an HTTP health endpoint (Caido's
forward proxy). We open the socket and immediately close it — the
handshake completing is enough to confirm readiness.
"""
deadline = asyncio.get_event_loop().time() + timeout
last_error: str | None = None
while asyncio.get_event_loop().time() < deadline:
try:
reader, writer = await asyncio.wait_for(
asyncio.open_connection(host, port),
timeout=poll_interval * 4,
)
except (TimeoutError, OSError) as e:
last_error = type(e).__name__
else:
writer.close()
# Some servers close hard immediately after accept; we only
# care that the connect itself succeeded.
with contextlib.suppress(OSError):
await writer.wait_closed()
del reader
return
await asyncio.sleep(poll_interval)
raise SandboxNotReadyError(
f"TCP probe of {host}:{port} did not connect within {timeout}s (last error: {last_error})",
)
+204
View File
@@ -0,0 +1,204 @@
"""Per-scan sandbox session lifecycle.
Replaces the legacy ``DockerRuntime`` (``strix/runtime/docker_runtime.py``)
with the SDK-native session model. One session per scan, reused across
every agent in that scan's tree.
The bundle returned by :func:`create_or_reuse` is what the per-agent
context dict reads from in ``run_config_factory.make_agent_context`` —
``client``, ``session``, ``tool_server_host_port``, ``caido_host_port``,
and ``bearer`` for authenticating to the in-container FastAPI tool server.
Cache strategy: a module-level dict keyed by ``scan_id``. The same scan
issuing multiple ``create_or_reuse`` calls (e.g., resume after a crash
on the host side) gets the same bundle back. ``cleanup`` is best-effort
— a leaked container is preferable to a stuck cleanup that prevents the
next scan from starting.
References:
- PLAYBOOK.md §3.3
- HARNESS_WIKI.md §6 (legacy Docker runtime)
"""
from __future__ import annotations
import logging
import secrets
import socket
from typing import TYPE_CHECKING, Any
import docker
from agents.sandbox.entries import LocalDir
from agents.sandbox.manifest import Environment, Manifest
from agents.sandbox.sandboxes.docker import DockerSandboxClientOptions
from strix.runtime.strix_docker_client import StrixDockerSandboxClient
from strix.sandbox.caido_capability import CaidoCapability
if TYPE_CHECKING:
from pathlib import Path
logger = logging.getLogger(__name__)
# In-container ports (must match the image's tool server + Caido sidecar
# binds). Defined here as a single source of truth for both the
# capability and the manifest env vars.
_CONTAINER_TOOL_SERVER_PORT = 48081
_CONTAINER_CAIDO_PORT = 48080
# Per-scan session cache. Module-level so a scan that bounces through
# multiple host-side processes (e.g., re-imports the module) doesn't
# spin up a second container — though in practice we expect one
# Strix process per scan.
_SESSION_CACHE: dict[str, dict[str, Any]] = {}
def _alloc_loopback_port() -> int:
"""Reserve a free 127.0.0.1 port via ephemeral socket bind.
Used only as a fallback when the SDK doesn't return a resolved
host port (older SDK versions before ``_resolve_exposed_port``
existed). Modern path uses the SDK's resolution.
"""
sock = socket.socket()
try:
sock.bind(("127.0.0.1", 0))
return int(sock.getsockname()[1])
finally:
sock.close()
async def create_or_reuse(
scan_id: str,
*,
image: str,
sources_path: Path,
execution_timeout: int = 120,
) -> dict[str, Any]:
"""Return the existing bundle for ``scan_id`` or create a new one.
Args:
scan_id: Caller-provided scan identifier (used as cache key).
image: Docker image tag (e.g. ``"strix-sandbox:0.1.13"``).
sources_path: Host directory mounted into the container's
``/workspace/sources`` so the agent can read user code.
execution_timeout: ``STRIX_SANDBOX_EXECUTION_TIMEOUT`` env var
inside the container — caps how long the in-container tool
server waits for a tool to finish before responding 504.
Defaults to 120s, matching the legacy harness.
Returns the bundle dict containing ``client``, ``session``,
``tool_server_host_port``, ``caido_host_port``, ``bearer``,
and ``capability`` (the live CaidoCapability instance).
"""
cached = _SESSION_CACHE.get(scan_id)
if cached is not None:
logger.info("Reusing existing sandbox session for scan %s", scan_id)
return cached
bearer = secrets.token_urlsafe(32)
capability = CaidoCapability()
# ``Manifest.environment`` is an ``Environment`` model — a bare dict
# is silently dropped by Pydantic, so wrap explicitly.
manifest = Manifest(
entries={"sources": LocalDir(src=sources_path)},
environment=Environment(
value={
"TOOL_SERVER_TOKEN": bearer,
"TOOL_SERVER_PORT": str(_CONTAINER_TOOL_SERVER_PORT),
"CAIDO_PORT": str(_CONTAINER_CAIDO_PORT),
"STRIX_SANDBOX_EXECUTION_TIMEOUT": str(execution_timeout),
"PYTHONUNBUFFERED": "1",
"HOST_GATEWAY": "host.docker.internal",
},
),
)
# The SDK's DockerSandboxClient requires a docker.DockerClient instance
# at construction time (since openai-agents 0.14.x). We use the
# caller's environment to find the daemon — same as the legacy
# DockerRuntime did via ``docker.from_env()``.
client = StrixDockerSandboxClient(docker.from_env())
options = DockerSandboxClientOptions(
image=image,
exposed_ports=(_CONTAINER_TOOL_SERVER_PORT, _CONTAINER_CAIDO_PORT),
)
logger.info(
"Creating sandbox session for scan %s (image=%s, exec_timeout=%ds)",
scan_id,
image,
execution_timeout,
)
session = await client.create(options=options, manifest=manifest)
# Resolve the host-side mapped ports the SDK assigned. The capability
# needs these *before* it binds, so its healthcheck task probes the
# right ports.
tool_server_endpoint = await session._resolve_exposed_port(
_CONTAINER_TOOL_SERVER_PORT,
)
caido_endpoint = await session._resolve_exposed_port(_CONTAINER_CAIDO_PORT)
capability.configure_host_ports(
tool_server_host_port=tool_server_endpoint.port,
caido_host_port=caido_endpoint.port,
)
# Bind the capability against the live session — this schedules the
# healthcheck task that on_agent_start awaits.
capability.bind(session)
bundle = {
"client": client,
"session": session,
"capability": capability,
"tool_server_host_port": tool_server_endpoint.port,
"caido_host_port": caido_endpoint.port,
"bearer": bearer,
}
_SESSION_CACHE[scan_id] = bundle
return bundle
async def cleanup(scan_id: str) -> None:
"""Tear down ``scan_id``'s container and drop its cache entry.
Best-effort: any error during ``client.delete`` is logged and
swallowed. We never want a cleanup failure to prevent the next
scan from starting; the worst case is a stranded container that
Docker's normal reaping will catch on next ``docker prune``.
"""
bundle = _SESSION_CACHE.pop(scan_id, None)
if bundle is None:
logger.debug("cleanup(%s): no cached session", scan_id)
return
capability = bundle.get("capability")
if isinstance(capability, CaidoCapability):
task = capability._healthcheck_task
if task is not None and not task.done():
task.cancel()
try:
await bundle["client"].delete(bundle["session"])
logger.info("Cleaned up sandbox session for scan %s", scan_id)
except Exception:
logger.exception(
"cleanup(%s): client.delete raised; container may need manual reaping",
scan_id,
)
def cached_scan_ids() -> list[str]:
"""Snapshot of currently-cached scan ids. Used by the TUI / CLI."""
return list(_SESSION_CACHE.keys())
def _reset_cache_for_tests() -> None:
"""Test helper — clears the module cache between unit tests."""
_SESSION_CACHE.clear()