feat(migration): phase 5 — root agent factory + entry point

Three new modules that wire Phases 0-4 into a runnable Strix scan:

- strix/agents/sdk_prompt.py: standalone Jinja-based system prompt
  renderer. Reuses the existing strix/agents/StrixAgent/system_prompt.
  jinja template (508 lines, the actual production prompt) so behavior
  parity with the legacy LLM._load_system_prompt is byte-identical.
  Skill resolution mirrors LLM._get_skills_to_load (caller skills →
  scan_modes/<mode> → whitebox pair, deduped). Fail-soft: template
  errors return empty string and log; agent construction must never
  blow up on prompt load.

- strix/agents/sdk_factory.py: build_strix_agent(name, skills, is_root)
  assembles an agents.Agent. Root carries finish_scan and stops there;
  child carries agent_finish and stops there (C4). Caido tools come
  from CaidoCapability automatically — we don't include them in
  _BASE_TOOLS to avoid double-registration when the SDK runtime merges
  capability tools. model=None so RunConfig drives the model alias
  through MultiProvider rather than the SDK default. make_child_factory
  returns a closure over scan-level config (scan_mode, is_whitebox,
  interactive, scope context) for ctx.context['agent_factory'] — the
  Phase 3 create_agent tool calls it with (name, skills) per child.

- strix/sdk_entry.py: run_strix_scan() — the top-level coroutine.
  Builds the bus, brings up (or reuses) a sandbox session via
  session_manager, builds the root Agent and the child factory, builds
  the per-agent context dict, registers the root in the bus, builds
  the RunConfig, calls Runner.run, and cleans up the session in a
  finally. Cancels descendants before re-raising any exception (C9).
  cleanup_on_exit toggle preserves the cached session for resume
  scenarios. _build_root_task and _build_scope_context preserve the
  legacy StrixAgent.execute_scan task formatting + scope context shape
  so the prompt template sees identical inputs.

Tests: 21 new tests (10 for factory + prompt, 11 for entry point).
Factory: root vs child tool list parity, finish_scan/agent_finish
placement, tool_use_behavior dict shape, Caido absence (capability-
provided), make_child_factory closure semantics. Entry point (all
mocked, no real Docker/LLM): wiring shape verification — context dict
carries every field downstream consumers read, session manager called
with correct scan_id, cleanup runs even on Runner.run failure,
cleanup skipped when disabled, scan_id auto-generation, scan-level
config (scan_mode, is_whitebox) flows into the factory. Task and scope
builders verified against the same shape as legacy.

Per-file ruff ignores added: TC002 on sdk_factory (Tool used at
runtime in _BASE_TOOLS tuple), TC003 + PLR0912 on sdk_entry (Path
runtime-imported; _build_root_task's per-target-type branches are
intentional and well-bounded).
This commit is contained in:
0xallam
2026-04-25 00:58:32 -07:00
parent 775a78487d
commit f9fcfd4edf
6 changed files with 1162 additions and 0 deletions
+221
View File
@@ -0,0 +1,221 @@
"""build_strix_agent — assemble an ``agents.Agent`` for root or child runs.
This is the keystone that links Phase 2's SDK function tools, Phase 3's
graph tools, Phase 4's CaidoCapability, and the rendered Jinja prompt
from :mod:`strix.agents.sdk_prompt` into a single ``agents.Agent``
instance ready for ``Runner.run``.
Two flavors:
- **Root** (``is_root=True``): the top-level scan agent. Carries
``finish_scan`` (terminates the scan), no ``agent_finish`` (that's
for subagents). ``tool_use_behavior`` stops on ``finish_scan`` so
the model can't accidentally keep talking after marking the scan
complete.
- **Child** (``is_root=False``): subagents spawned by the
``create_agent`` graph tool. Carries ``agent_finish``, no
``finish_scan``. ``tool_use_behavior`` stops on ``agent_finish``
(C4 — without this, the SDK loop would keep going to ``max_turns``
even after the child reported back to its parent).
Caido tools come from ``CaidoCapability.tools()`` automatically via
the SDK's capability merge — we don't include them here. Skills are
injected via the prompt; the model can also load more at runtime via
the ``load_skill`` tool.
References:
- PLAYBOOK.md §4.3 (graph tool wiring)
- AUDIT.md §2.4 (C4 — stop_at_tool_names is required for subagents)
"""
from __future__ import annotations
import logging
from typing import Any
from agents import Agent
from agents.agent import StopAtTools
from agents.tool import Tool
from strix.agents.sdk_prompt import render_system_prompt
from strix.tools.agents_graph.agents_graph_sdk_tools import (
agent_finish,
agent_status,
create_agent,
send_message_to_agent,
view_agent_graph,
wait_for_message,
)
from strix.tools.browser.browser_sdk_tool import browser_action
from strix.tools.file_edit.file_edit_sdk_tools import (
list_files,
search_files,
str_replace_editor,
)
from strix.tools.finish.finish_sdk_tool import finish_scan
from strix.tools.load_skill.load_skill_sdk_tool import load_skill
from strix.tools.notes.notes_sdk_tools import (
create_note,
delete_note,
get_note,
list_notes,
update_note,
)
from strix.tools.python.python_sdk_tool import python_action
from strix.tools.reporting.reporting_sdk_tools import create_vulnerability_report
from strix.tools.terminal.terminal_sdk_tool import terminal_execute
from strix.tools.thinking.thinking_sdk_tools import think
from strix.tools.todo.todo_sdk_tools import (
create_todo,
delete_todo,
list_todos,
mark_todo_done,
mark_todo_pending,
update_todo,
)
from strix.tools.web_search.web_search_sdk_tool import web_search
logger = logging.getLogger(__name__)
# Tools every Strix agent has, root or child. The Caido proxy tools
# (list_requests, view_request, send_request, ...) are NOT here —
# CaidoCapability.tools() returns them and the SDK merges them in.
_BASE_TOOLS: tuple[Tool, ...] = (
# Thinking + planning
think,
# Per-agent todos
create_todo,
list_todos,
update_todo,
mark_todo_done,
mark_todo_pending,
delete_todo,
# Shared notes (per-run JSONL store)
create_note,
list_notes,
get_note,
update_note,
delete_note,
# Web search (only registered if PERPLEXITY_API_KEY is set; the
# tool itself returns a structured error when not configured, so
# always exposing it is safe)
web_search,
# File edit (sandbox-bound)
str_replace_editor,
list_files,
search_files,
# Reporting
create_vulnerability_report,
# Skill loading
load_skill,
# Sandbox primitives
browser_action,
terminal_execute,
python_action,
# Multi-agent graph tools (the bus is in ctx.context)
view_agent_graph,
agent_status,
send_message_to_agent,
wait_for_message,
create_agent,
)
def build_strix_agent(
*,
name: str = "strix",
skills: list[str] | None = None,
is_root: bool,
scan_mode: str = "deep",
is_whitebox: bool = False,
interactive: bool = False,
system_prompt_context: dict[str, Any] | None = None,
) -> Agent[Any]:
"""Build an ``agents.Agent`` configured for either root or child use.
Args:
name: Agent name. Surfaces in traces and the bus's ``names`` map.
Defaults to ``"strix"`` for the root; create_agent passes
distinct names per child.
skills: Skills to preload into the system prompt. The agent can
also load more at runtime via the ``load_skill`` tool.
is_root: Selects the tool list and ``tool_use_behavior``.
Root carries ``finish_scan`` and stops there; child carries
``agent_finish`` and stops there.
scan_mode: ``"deep"`` etc.; routes the scan-mode skill section
of the prompt template.
is_whitebox: Whitebox source-aware mode toggle. Adds two extra
skills to the prompt and gates whitebox-only behavior in
the create_agent / wiki integration.
interactive: Renders the interactive-mode communication block
in the system prompt.
system_prompt_context: Free-form dict the prompt template
renders into the ``system_prompt_context`` variable —
today carries the scan scope / authorization block.
Returns the ``Agent`` instance with ``model=None`` so the
``RunConfig.model`` (built by ``make_run_config``) drives provider
selection. ``agents.Agent`` is generic on context type; we let
the caller's ``Runner.run(context=...)`` typing determine that.
"""
instructions = render_system_prompt(
skills=skills,
scan_mode=scan_mode,
is_whitebox=is_whitebox,
interactive=interactive,
system_prompt_context=system_prompt_context,
)
# Tool list + termination tool depend on is_root. The tuple-then-
# list dance keeps _BASE_TOOLS immutable so concurrent agent builds
# can't accidentally mutate each other's tool list.
if is_root:
tools: list[Tool] = [*_BASE_TOOLS, finish_scan]
stop_at = ("finish_scan",)
else:
tools = [*_BASE_TOOLS, agent_finish]
stop_at = ("agent_finish",)
return Agent(
name=name,
instructions=instructions,
tools=tools,
tool_use_behavior=StopAtTools(stop_at_tool_names=list(stop_at)),
# model=None so ``RunConfig.model`` (e.g. ``strix/claude-sonnet-4.6``)
# routes through MultiProvider rather than the SDK's default.
model=None,
)
def make_child_factory(
*,
scan_mode: str = "deep",
is_whitebox: bool = False,
interactive: bool = False,
system_prompt_context: dict[str, Any] | None = None,
) -> Any:
"""Return a callable suitable for ``ctx.context['agent_factory']``.
The Phase 3 ``create_agent`` graph tool reads
``ctx.context['agent_factory']`` and calls it with ``name=`` and
``skills=`` to build a child Agent. We snapshot the run-level
arguments (scan_mode, is_whitebox, etc.) into a closure so each
child inherits the right scan-level configuration without the
create_agent tool having to know about them.
"""
def _factory(*, name: str, skills: list[str]) -> Agent[Any]:
return build_strix_agent(
name=name,
skills=skills,
is_root=False,
scan_mode=scan_mode,
is_whitebox=is_whitebox,
interactive=interactive,
system_prompt_context=system_prompt_context,
)
return _factory
+130
View File
@@ -0,0 +1,130 @@
"""Standalone Jinja-based system-prompt renderer for SDK agents.
The legacy ``LLM._load_system_prompt`` couples prompt rendering to the
LLM client class. The SDK migration owns the model client through
``MultiProvider`` instead, so we extract the rendering logic into a
plain function that the SDK agent factory can call without pulling in
the legacy ``LLM`` instance.
Reuses the existing Jinja template at
``strix/agents/StrixAgent/system_prompt.jinja`` (508 lines, expanding
into the multi-section prompt with skills, tools, scan modes, etc.) so
behavior parity is preserved verbatim — only the call site changes.
References:
- HARNESS_WIKI.md §4.1 (system prompt assembly)
- PLAYBOOK.md §4 (per-tool migration contracts)
"""
from __future__ import annotations
import logging
from typing import Any
from jinja2 import Environment, FileSystemLoader, select_autoescape
from strix.skills import load_skills
from strix.tools import get_tools_prompt
from strix.utils.resource_paths import get_strix_resource_path
logger = logging.getLogger(__name__)
# Hard-coded to the StrixAgent template since it's the only agent type
# under the SDK migration. The legacy harness supported multiple agent
# names but in practice only StrixAgent ships.
_AGENT_NAME = "StrixAgent"
def _resolve_skills(
*,
requested: list[str] | None,
scan_mode: str = "deep",
is_whitebox: bool = False,
) -> list[str]:
"""Build the deduped, ordered skills list for the prompt render.
Mirrors :py:meth:`LLM._get_skills_to_load` exactly so the rendered
prompt is byte-identical to the legacy path:
1. Whatever the caller asked for, in order.
2. ``scan_modes/<mode>`` (always).
3. Whitebox-specific skills if applicable.
"""
ordered: list[str] = list(requested or [])
ordered.append(f"scan_modes/{scan_mode}")
if is_whitebox:
ordered.append("coordination/source_aware_whitebox")
ordered.append("custom/source_aware_sast")
deduped: list[str] = []
seen: set[str] = set()
for skill in ordered:
if skill and skill not in seen:
deduped.append(skill)
seen.add(skill)
return deduped
def render_system_prompt(
*,
skills: list[str] | None = None,
scan_mode: str = "deep",
is_whitebox: bool = False,
interactive: bool = False,
system_prompt_context: dict[str, Any] | None = None,
) -> str:
"""Render the StrixAgent system prompt.
Args:
skills: Skills the caller wants preloaded into the prompt
context (the agent can also load more at runtime via the
``load_skill`` tool).
scan_mode: ``"deep" | "fast" | ...``. Maps to ``scan_modes/<mode>``
skill.
is_whitebox: When True, the source-aware whitebox skill stack
is loaded too.
interactive: When True, the prompt renders the interactive-mode
communication rules block.
system_prompt_context: Free-form dict that the template's
``system_prompt_context`` variable receives — used today for
the scan-scope authorization block from
:py:meth:`StrixAgent._build_system_scope_context`.
Returns the rendered prompt string. If anything goes wrong (template
missing, render failure), returns an empty string and logs — same
fail-soft posture as the legacy method, because a missing prompt is
survivable but a hard failure during agent construction is not.
"""
try:
prompt_dir = get_strix_resource_path("agents", _AGENT_NAME)
skills_dir = get_strix_resource_path("skills")
env = Environment(
loader=FileSystemLoader([prompt_dir, skills_dir]),
autoescape=select_autoescape(
enabled_extensions=(),
default_for_string=False,
),
)
skills_to_load = _resolve_skills(
requested=skills,
scan_mode=scan_mode,
is_whitebox=is_whitebox,
)
skill_content = load_skills(skills_to_load)
env.globals["get_skill"] = lambda name: skill_content.get(name, "")
rendered = env.get_template("system_prompt.jinja").render(
get_tools_prompt=get_tools_prompt,
loaded_skill_names=list(skill_content.keys()),
interactive=interactive,
system_prompt_context=system_prompt_context or {},
**skill_content,
)
except Exception:
logger.exception("render_system_prompt failed; returning empty prompt")
return ""
else:
return str(rendered)