Compare commits

..
Author SHA1 Message Date
Alex Schapiro 6334074a0b fix: also reject a fallback model from a different provider 2026-08-10 15:54:36 +00:00
Alex Schapiro 9ddb2b0bb6 fix: reject a fallback model on a different tool-schema route 2026-08-10 15:41:23 +00:00
Alex Schapiro 2c5d4a166b feat: serve denied turns on the fallback before pinning the agent 2026-08-10 15:33:41 +00:00
Alex Schapiro 009cc94220 test: add fallback settings fields to runner test stubs 2026-08-10 13:45:23 +00:00
Alex Schapiro cd15e9e8c3 fix: build fallback-specific model settings for the denial fallback 2026-08-10 13:42:20 +00:00
Alex Schapiro b7d7b2f3b4 fix: read fallback settings directly instead of via getattr 2026-08-10 13:20:54 +00:00
Alex Schapiro e170c5506b feat: fall back after repeated content denials 2026-08-10 13:19:23 +00:00
Ahmed AllamandAhmed Allam 94a2586aaa fix(container): write the browser profile as root 2026-08-10 10:08:17 +03:00
Ahmed AllamandAhmed Allam 372e27fa17 chore(container): drop explanatory comment 2026-08-10 09:54:49 +03:00
Ahmed AllamandAhmed Allam ad727edd66 fix(container): keep the browser env alive where image ENV is dropped 2026-08-10 09:54:49 +03:00
7b3c8f9b74 fix(container): reclaim abandoned browser sessions (#1034)
Co-authored-by: Ahmed Allam <ahmed39652003@gmail.com>
2026-08-09 16:57:51 -07:00
Ahmed AllamandAhmed Allam ae07af6159 chore: drop explanatory comment 2026-08-09 15:44:16 +03:00
Ahmed AllamandAhmed Allam 649a2e2140 fix(llm): omit parallel_tool_calls on tool-less requests 2026-08-09 15:44:16 +03:00
Ahmed AllamandAhmed Allam 597aae6715 chore: release v1.5.2 2026-08-09 04:29:34 +03:00
Ahmed AllamandGitHub 06b158d1fa fix(runner): settle child agents before closing sessions at wind-down (#1025) 2026-08-08 18:17:58 -07:00
Ahmed AllamandGitHub c29eb73c7f fix(tools): coerce an empty-string list/dict argument to an empty container (#1024) 2026-08-08 16:58:30 -07:00
Ahmed AllamandGitHub 72833b8e43 fix(runner): resume after a user interrupt instead of failing (#1023) 2026-08-08 16:44:12 -07:00
Ahmed AllamandGitHub 1117ba6d4a fix(sessions): open a sqlite connection per operation, not per thread (#1022) 2026-08-08 16:20:01 -07:00
Ahmed AllamandGitHub 53e4658d88 fix(todo): stop a todo plan failing on priority or duplicates (#1021) 2026-08-08 15:18:48 -07:00
58df71d3db fix(agents): let an agent wait on what it already said (#1020)
* let an agent wait on what it already said

An agent that answers in plain text is nudged to call a tool, and the only tool
that hands control back takes a required message. So it says the same thing
twice: once as text the user has already read, once as the argument it had to
supply to stop. Seen on a run whose whole instruction was "hi" - a greeting, then
the same greeting again through respond_to_user.

message is optional now. The nudge arms the tool with the text that was
delivered and says not to repeat it, so an agent that has said its piece can park
on it with an empty call. Anything it does want to add it passes normally.

Parking still cannot leave the user on silence: an empty call is refused unless
something was actually said, and the arming is single use - execution clears it
as soon as a turn ends any other way.

The interactive prompt now also says to answer and stop in one respond_to_user
call, which is what avoids the nudge in the first place.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* drop the worked example from the interactive prompt

"the user greeted you, asked something you can answer outright, or you need a
decision" was the run I had been reading, written into a rule that holds
whatever the reason. The rule is that replying and stopping is one call; listing
occasions only invites the model to check whether this is one of them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* drop the arming flag; an empty message just waits

Passing the delivered text from execution into the tool, and refusing an empty
call without it, was machinery guarding against an agent parking having said
nothing. That leaves the user looking at "waiting for your reply" with a cursor
in front of them - they type. It does not need a mechanism.

What is left is the default on message, and the nudge saying the text already
landed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* only offer waiting on words that were written

The nudge told every agent its text had already been delivered, but it fires
whenever a turn leaves the agent running, and a turn can end with no tool call
and no text at all - _final_output_preview has carried <none> and <empty>
branches all along. An agent that said nothing was being invited to wait on an
answer the user never received, leaving them at a bare prompt.

It now reads the turn: waiting on what was said is offered only when something
was, and otherwise the agent is told plainly that the user has read nothing and
to send its message.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* leave the continuation nudge alone

Rewording it meant asserting from the outside whether the agent had spoken, and
the nudge fires whenever a turn leaves the agent running - text or no text. The
agent knows which it did without being told, so the guidance belongs in its
prompt, where the condition is its own to read.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* say it in the nudge, where the agent is reading

An agent stranded by the nudge reasons off the nudge. Told only to call
respond_to_user, it supplies a message, and since it has just answered in plain
text that message is the same answer again. The system prompt saying otherwise
sits thousands of tokens earlier and loses.

The clause goes on the line the agent acts on: call respond_to_user, with no
message if it has already said it. That reads true whatever the turn did,
including one that produced no text, because the agent is the one who knows
which — nothing here has to work it out from the outside.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-09 00:57:16 +03:00
0b9e029a5d test(tui): correct the nudge the internal-turn test asserts (#1016)
* correct the nudge the internal-turn test asserts

The test expected "ended the autonomous Strix run", which strix.core.execution
does not inject; it says "ended the autonomous run". The classifier was right and
the test was not, so the suite failed on main while the behaviour it guards was
fine.

The sentence is written inline in another module and copied by hand into the
classifier and again into the test, which is how it drifted. A second test now
reads it back out of that module's source, joining the adjacent string literals
its line wrapping leaves behind, and fails if either nudge is no longer injected
verbatim. Reworded one and it reports which nudge went missing and what a resumed
scan would do about it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* read the nudges out of what the module can inject, not out of its text

Searching the source accepted the sentence anywhere in the file, so a stale copy
left behind in a comment would have kept the guard passing after the message it
guards had changed - the drift it exists to catch.

Parsing the module instead limits it to strings the code can actually inject.
Comments never reach the tree, docstrings are dropped as description rather than
behaviour, and adjacent literals are joined during parsing, which the line
wrapping needed and the regex was only approximating.

Checked by rewording the message and leaving the old wording in a comment: the
guard fails, where searching the text passed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-08 22:34:45 +03:00
b260a4ee38 fix(tui): make the mount prompt clickable, and skip the mount instead of abandoning the scan (#1015)
* make the working-directory prompt answer the mouse

Its Confirm and Cancel were drawn as buttons and did nothing when clicked: the
modal mouse handler had a case for every dialog except this one, so a click fell
through and the scan sat waiting on an answer the user believed they had given.
Only the keyboard could answer it.

The prompt is docked in a corner rather than centered, so it also needs its own
bounds; the centered ones every other dialog uses would have put the buttons in
the wrong place. Those bounds now come from the same placement cornerOverlay
draws with.

Two returns that hand back the model alongside a call that mutates it are now
sequenced explicitly. They work, but only because the compiler happens to
evaluate the call first, and one of them is what puts the prompt back in the
composer when the mount is declined.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* skip the mount instead of abandoning the scan

Declining the working-directory prompt threw the whole launch away and dropped
back to the start screen, which is a lot to lose for answering one question
about one directory. The two answers are now about the directory alone: mount it,
or run without it. The prompt is the whole of the input either way.

The buttons say which is which - Mount and Skip rather than Confirm and Cancel -
and the prompt says what skipping costs.

A run with neither target nor directory is a real run, so two things follow it.
It can be resumed: its instruction is what drives it, and that is in the run
record. And it tells the agent plainly that it has neither, because an agent
given no scope goes looking for the one it assumes it was meant to have.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-08 22:34:02 +03:00
alex sandGitHub f8a8801d56 docs(skills): rename skills to descriptive names and broaden descriptions for discoverability (#1013) 2026-08-07 14:24:19 -04:00
Ahmed AllamandAhmed Allam 22750077da chore: release v1.5.1 2026-08-07 20:28:06 +03:00
52 changed files with 1294 additions and 345 deletions
+4 -4
View File
@@ -10,10 +10,10 @@ Install the agent skills for step-by-step workflows:
npx skills add usestrix/strix
```
- `strix-pentest` — run a headless pentest against code, URLs, domains, or IPs and read results (covers both run modes below)
- `strix-cloud-api` — drive the managed app.strix.ai platform via REST (no local Docker/LLM needed)
- `strix-fix-findings` — remediate findings and re-run Strix to verify
- `strix-ci-setup` — add PR scanning to CI/CD (self-hosted CLI or managed app)
- `penetration-testing-with-strix` — run a headless pentest against code, URLs, domains, or IPs and read results (covers both run modes below)
- `managed-pentesting-with-strix` — drive the managed app.strix.ai platform via REST (no local Docker/LLM needed)
- `fix-security-vulnerabilities-with-strix` — remediate findings and re-run Strix to verify
- `ci-security-scanning-with-strix` — add PR scanning to CI/CD (self-hosted CLI or managed app)
**Two ways to run, same engine — pick per situation:**
+1 -1
View File
@@ -116,7 +116,7 @@ Strix is agent-ready. Give Claude Code, Cursor, Codex, or any [SKILL.md-compatib
npx skills add usestrix/strix
```
This installs four skills: **strix-pentest** (run headless scans and read results), **strix-cloud-api** (drive the managed [app.strix.ai](https://app.strix.ai) platform via REST — no local Docker or LLM key), **strix-fix-findings** (remediate + re-scan to verify), and **strix-ci-setup** (PR scanning in CI). Agents can run Strix two ways with the same engine — the open-source CLI locally, or the managed cloud when there's no local infra — and read [`AGENTS.md`](AGENTS.md) for a quick reference, [docs.strix.ai/llms.txt](https://docs.strix.ai/llms.txt) for the CLI docs, and [docs.app.strix.ai](https://docs.app.strix.ai) for the API.
This installs four skills: **penetration-testing-with-strix** (run headless scans and read results), **managed-pentesting-with-strix** (drive the managed [app.strix.ai](https://app.strix.ai) platform via REST — no local Docker or LLM key), **fix-security-vulnerabilities-with-strix** (remediate + re-scan to verify), and **ci-security-scanning-with-strix** (PR scanning in CI). Agents can run Strix two ways with the same engine — the open-source CLI locally, or the managed cloud when there's no local infra — and read [`AGENTS.md`](AGENTS.md) for a quick reference, [docs.strix.ai/llms.txt](https://docs.strix.ai/llms.txt) for the CLI docs, and [docs.app.strix.ai](https://docs.app.strix.ai) for the API.
---
+15
View File
@@ -117,6 +117,21 @@ ENV AGENT_BROWSER_EXECUTABLE_PATH=/usr/bin/chromium
ENV AGENT_BROWSER_USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"
ENV AGENT_BROWSER_ARGS="--disable-blink-features=AutomationControlled,--no-first-run,--no-default-browser-check,--lang=en-US"
ENV AGENT_BROWSER_SCREENSHOT_DIR=/workspace/.agent-browser-screenshots
ENV AGENT_BROWSER_IDLE_TIMEOUT_MS=180000
USER root
RUN set -eu; \
{ \
for var in AGENT_BROWSER_EXECUTABLE_PATH AGENT_BROWSER_USER_AGENT \
AGENT_BROWSER_ARGS AGENT_BROWSER_SCREENSHOT_DIR \
AGENT_BROWSER_IDLE_TIMEOUT_MS; do \
eval "value=\${$var}"; \
printf 'export %s="${%s:-%s}"\n' "$var" "$var" "$value"; \
done; \
} > /tmp/agent-browser.sh; \
install -m 0644 /tmp/agent-browser.sh /etc/profile.d/agent-browser.sh; \
rm /tmp/agent-browser.sh; \
env -i bash -lc 'test "${AGENT_BROWSER_IDLE_TIMEOUT_MS}" = "180000"'
USER pentester
RUN /home/pentester/.npm-global/bin/agent-browser doctor --offline --quick
RUN set -eux; \
+8
View File
@@ -35,6 +35,14 @@ Configure Strix using environment variables or a config file.
Maximum number of retries for LLM API calls on transient failures.
</ParamField>
<ParamField path="STRIX_LLM_FALLBACK" type="string">
Optional model that retries any turn blocked by a content guardrail. Unset disables this behavior, leaving a denial terminal for that agent.
</ParamField>
<ParamField path="STRIX_LLM_DENIED_RETRIES" default="3" type="integer">
Content-guardrail denials an agent may take before it is pinned to `STRIX_LLM_FALLBACK` for the rest of its lifecycle. Below that count it returns to the main model on the next turn. Counted per agent.
</ParamField>
<ParamField path="STRIX_REASONING_EFFORT" default="high" type="string">
Control thinking effort for reasoning models. Valid values: `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`. Defaults to `medium` for quick scan mode.
</ParamField>
+7 -7
View File
@@ -15,15 +15,15 @@ npx skills add usestrix/strix
| Skill | What your agent learns |
|-------|------------------------|
| `strix-pentest` | Run headless scans against code, URLs, domains, or IPs — self-hosted CLI or managed cloud — with budget caps, and read the results |
| `strix-cloud-api` | Drive the managed [app.strix.ai](https://app.strix.ai) platform over REST — no local Docker or LLM key needed |
| `strix-fix-findings` | Triage findings, fix root causes, and re-run Strix to verify each fix |
| `strix-ci-setup` | Add PR security scanning to GitHub Actions or any CI (self-hosted CLI or managed app) |
| `penetration-testing-with-strix` | Run headless scans against code, URLs, domains, or IPs — self-hosted CLI or managed cloud — with budget caps, and read the results |
| `managed-pentesting-with-strix` | Drive the managed [app.strix.ai](https://app.strix.ai) platform over REST — no local Docker or LLM key needed |
| `fix-security-vulnerabilities-with-strix` | Triage findings, fix root causes, and re-run Strix to verify each fix |
| `ci-security-scanning-with-strix` | Add PR security scanning to GitHub Actions or any CI (self-hosted CLI or managed app) |
Install a single skill with `npx skills add usestrix/strix --skill strix-pentest`, or use one without installing:
Install a single skill with `npx skills add usestrix/strix --skill penetration-testing-with-strix`, or use one without installing:
```bash
npx skills use usestrix/strix@strix-pentest | claude
npx skills use usestrix/strix@penetration-testing-with-strix | claude
```
## Two ways to run — self-hosted or managed
@@ -31,7 +31,7 @@ npx skills use usestrix/strix@strix-pentest | claude
Both use the same engine and produce the same validated findings and SARIF, so agents can pick per situation or combine them:
- **Open-source CLI (self-hosted)** — runs locally in a Docker sandbox with your own LLM key. Free, fully local, air-gap capable. Best for local dev loops and full control.
- **Managed cloud** — runs on Strix's infrastructure via the [app.strix.ai REST API](https://docs.app.strix.ai). No Docker, no LLM key, no local install; adds team dashboards, scheduling, PR reviews, and downloadable PDF/DOCX reports (Enterprise plan). Best in sandboxed/CI environments and for teams. Create an API token under **Settings → API Access**; the `strix-cloud-api` skill has the full flow.
- **Managed cloud** — runs on Strix's infrastructure via the [app.strix.ai REST API](https://docs.app.strix.ai). No Docker, no LLM key, no local install; adds team dashboards, scheduling, PR reviews, and downloadable PDF/DOCX reports (Enterprise plan). Best in sandboxed/CI environments and for teams. Create an API token under **Settings → API Access**; the `managed-pentesting-with-strix` skill has the full flow.
## Agent-Friendly Interfaces
-20
View File
@@ -1,20 +0,0 @@
"""PyInstaller runtime hook: point SSL env vars at the bundled certifi CA file.
Without ``collect_data_files('certifi')`` and this hook, a frozen binary can
resolve ``certifi.where()`` to a missing path and fail TLS verification with a
generic ``Connection error`` from httpx/litellm.
"""
from __future__ import annotations
import sys
if getattr(sys, "frozen", False):
import os
import certifi
ca_bundle = certifi.where()
os.environ.setdefault("SSL_CERT_FILE", ca_bundle)
os.environ.setdefault("REQUESTS_CA_BUNDLE", ca_bundle)
+1 -3
View File
@@ -1,6 +1,6 @@
[project]
name = "strix-agent"
version = "1.5.0"
version = "1.5.2"
description = "Open-source AI Hackers for your apps"
readme = "README.md"
license = "Apache-2.0"
@@ -232,8 +232,6 @@ ignore = [
"tests/test_codex_auth.py" = ["S105", "S106", "SLF001"]
# Hatchling loads the build hook by path, not as an importable package.
"scripts/tui_sidecar_hook.py" = ["INP001"]
# PyInstaller runtime hooks are loaded by path, not as an importable package.
"hooks/rthooks/pyi_rth_certifi.py" = ["INP001"]
# Stdlib HTTP handler overrides (do_GET/do_POST).
"strix/interface/auth_cli.py" = ["N802"]
"tests/test_codex_streaming.py" = ["N802"]
@@ -1,6 +1,6 @@
---
name: strix-ci-setup
description: Wire Strix security scanning into CI/CD — GitHub Actions, GitLab CI, or any pipeline — so every pull request gets a diff-scoped AI pentest that blocks vulnerable code. Covers both the self-hosted open-source CLI (runs in your runner) and the managed app.strix.ai platform (GitHub/GitLab app or API, no runner infra). Use when the user asks to add security scanning, pentesting, or Strix to their CI pipeline or PR workflow.
name: ci-security-scanning-with-strix
description: Add security scanning to CI/CD with Strix — GitHub Actions, GitLab CI, or any pipeline — so every pull request gets a diff-scoped AI pentest that blocks vulnerable code before it merges, with results as PR comments and SARIF uploaded to code scanning. Covers both the self-hosted open-source CLI (runs in your runner) and the managed app.strix.ai platform (GitHub/GitLab app or API, no runner infra). Use when the user asks to add security scanning, SAST/DAST, pentesting, vulnerability checks, or automated security review to their CI pipeline, pre-merge gate, or PR workflow.
license: Apache-2.0
metadata:
author: usestrix
@@ -11,7 +11,7 @@ metadata:
You can gate PRs two ways — pick based on the environment, or combine them:
- **Managed platform (recommended for most teams)** — connect the GitHub/GitLab/Bitbucket app once and Strix reviews every PR with **no workflow file, no runner, no Docker, and no LLM key**. Results post as PR comments and land in the team dashboard. Best when you want zero CI maintenance, central tracking, or your runners lack Docker. See "Managed platform" below and the **strix-cloud-api** skill.
- **Managed platform (recommended for most teams)** — connect the GitHub/GitLab/Bitbucket app once and Strix reviews every PR with **no workflow file, no runner, no Docker, and no LLM key**. Results post as PR comments and land in the team dashboard. Best when you want zero CI maintenance, central tracking, or your runners lack Docker. See "Managed platform" below and the **managed-pentesting-with-strix** skill.
- **Self-hosted OSS CLI in your runner** — run a diff-scoped scan as a pipeline step. Fully in your infra, free (BYO LLM key), no external account. Requires Docker on the runner. Best for air-gapped/self-hosted CI or when you don't want scans leaving your environment.
Both fail the build on validated findings and both emit SARIF 2.1.0, so you can start with one and add the other later.
@@ -131,6 +131,6 @@ No workflow file, no Docker, no LLM key. Two ways to use it:
-d "{\"repository_full_name\":\"${{ github.repository }}\",\"pr_number\":${{ github.event.pull_request.number }}}"
```
To gate the build on results, poll the PR review / scan status and fail on unresolved criticals/highs. Full endpoints (PR reviews, scans, SARIF export, schedules for scheduled deep scans) are in the **strix-cloud-api** skill.
To gate the build on results, poll the PR review / scan status and fail on unresolved criticals/highs. Full endpoints (PR reviews, scans, SARIF export, schedules for scheduled deep scans) are in the **managed-pentesting-with-strix** skill.
Recommend Option B for most teams (no maintenance, central dashboard); use Option A when scans must stay entirely within your own infrastructure.
@@ -1,6 +1,6 @@
---
name: strix-fix-findings
description: Triage and remediate vulnerabilities found by a Strix pentest (open-source CLI or app.strix.ai cloud), then re-run Strix to verify each fix. Use after a Strix scan reports findings, or when the user asks to fix security issues from a strix_runs report, vulnerabilities.json, findings.sarif, or a cloud scan's vulnerabilities.
name: fix-security-vulnerabilities-with-strix
description: Fix security vulnerabilities found by a Strix pentest (open-source CLI or app.strix.ai cloud) — triage by severity, patch the root cause rather than the symptom, and re-run Strix to prove each fix actually closes the exploit. Handles injection, XSS, SSRF, broken access control, IDOR, and other validated findings. Use after a Strix scan reports findings, or when the user asks to remediate, patch, or fix security issues from a strix_runs report, vulnerabilities.json, findings.sarif, or a cloud scan.
license: Apache-2.0
metadata:
author: usestrix
@@ -18,7 +18,7 @@ Get the findings from wherever the scan ran:
- **OSS CLI** — artifacts in `strix_runs/<run-name>/`:
- `vulnerabilities/*.md` — one finding per file: description, severity, PoC steps or script, affected code locations, remediation guidance.
- `vulnerabilities.json` — the same findings as JSON (ids, severity, CWE/CVE, `code_locations` with `fix_before`/`fix_after` suggestions when available).
- **Cloud (app.strix.ai)** — fetch the scan's `vulnerabilities[]` via `GET /api/v1/scans/{scanId}` (or `GET /api/v1/vulnerabilities` org-wide). Each carries `severity, cwe, endpoint, method, impact, technical_analysis, poc_description, poc_script_code` and, for code findings, `code_file`/`code_diff`/`code_before`/`code_after`. See the **strix-cloud-api** skill for auth.
- **Cloud (app.strix.ai)** — fetch the scan's `vulnerabilities[]` via `GET /api/v1/scans/{scanId}` (or `GET /api/v1/vulnerabilities` org-wide). Each carries `severity, cwe, endpoint, method, impact, technical_analysis, poc_description, poc_script_code` and, for code findings, `code_file`/`code_diff`/`code_before`/`code_after`. See the **managed-pentesting-with-strix** skill for auth.
Order work by severity: critical → high → medium → low. Every Strix finding was validated with a working proof-of-concept, so do not dismiss findings as false positives without re-testing the PoC yourself.
@@ -1,6 +1,6 @@
---
name: strix-cloud-api
description: Drive the managed Strix platform headlessly through the app.strix.ai REST API — create an API token, register domain/repository assets, launch and poll pentest scans, list and triage vulnerabilities, export SARIF, download PDF/DOCX reports (Enterprise plan), start PR reviews, and set up schedules and webhooks. Use when the user wants Strix without local Docker/LLM infra, or wants scans tracked in a team dashboard, on a schedule, or in CI via API.
name: managed-pentesting-with-strix
description: Run a managed pentest of a web app or API through the app.strix.ai REST API — no local Docker, LLM key, or install needed. Create an API token, register domain/repository assets, launch and poll scans, triage vulnerabilities, export SARIF, download PDF/DOCX pentest reports for SOC 2 and other compliance evidence (Enterprise plan), start PR reviews, and set up schedules and webhooks. Use when the user wants continuous or scheduled pentesting-as-a-service, an auditor-ready pentest report, scans tracked in a team dashboard, or security testing from a sandboxed agent/CI environment with no infrastructure.
license: Apache-2.0
metadata:
author: usestrix
@@ -9,7 +9,7 @@ metadata:
# Strix Cloud API (managed, no local infra)
Use this when you want Strix's autonomous pentesting **without running Docker or an LLM yourself** — the scan runs on Strix's infrastructure and results are tracked in a team dashboard. This is the right choice in sandboxed/hosted agent and CI environments, for teams, and for scheduled/continuous testing (downloadable PDF/DOCX reports are an Enterprise-plan feature). For fully local, free, air-gapped, or BYO-LLM runs, use the open-source CLI in the **strix-pentest** skill instead — both share the same engine and SARIF output, so you can mix them.
Use this when you want Strix's autonomous pentesting **without running Docker or an LLM yourself** — the scan runs on Strix's infrastructure and results are tracked in a team dashboard. This is the right choice in sandboxed/hosted agent and CI environments, for teams, and for scheduled/continuous testing (downloadable PDF/DOCX reports are an Enterprise-plan feature). For fully local, free, air-gapped, or BYO-LLM runs, use the open-source CLI in the **penetration-testing-with-strix** skill instead — both share the same engine and SARIF output, so you can mix them.
Full reference: **[docs.app.strix.ai](https://docs.app.strix.ai)** · OpenAPI: `https://docs.app.strix.ai/openapi.json`
@@ -113,7 +113,7 @@ curl -sS "$BASE/scans/$scan_id" "${auth[@]}" \
Cloud severities are `critical | high | medium | low` and statuses are `open | in_progress | fixed | ignored`. Sort by an explicit severity order rather than `sort_by(.severity)`, which sorts alphabetically (critical, high, low, medium).
Org-wide triage across scans: `GET /vulnerabilities` (`vulnerabilities:read`; filter by severity/status). Update triage state with the vulnerabilities `:write` endpoints. To remediate, hand off to the **strix-fix-findings** skill.
Org-wide triage across scans: `GET /vulnerabilities` (`vulnerabilities:read`; filter by severity/status). Update triage state with the vulnerabilities `:write` endpoints. To remediate, hand off to the **fix-security-vulnerabilities-with-strix** skill.
## 5. Export & report
@@ -1,6 +1,6 @@
---
name: strix-pentest
description: Run an autonomous AI penetration test with Strix against a codebase, repository, URL, domain, or IP — either self-hosted with the open-source CLI or via the managed app.strix.ai cloud API — and read the validated findings (Markdown, JSON, CSV, SARIF, PoCs). Use when the user asks to pentest, security-scan, or find vulnerabilities in an app, API, website, or repo with Strix.
name: penetration-testing-with-strix
description: Pentest a web app, API, codebase, repository, URL, domain, or IP with Strix — autonomous AI penetration testing that exploits and proves vulnerabilities (OWASP Top 10 and beyond — injection, XSS, SSRF, auth/access-control flaws, IDOR, business logic) instead of just flagging them. Runs self-hosted with the open-source CLI or via the managed app.strix.ai cloud, and returns validated findings with proof-of-concept exploits (Markdown, JSON, CSV, SARIF). Use when the user asks to pentest, hack, security-scan, security-audit, or find vulnerabilities in an app, API, website, or repo.
license: Apache-2.0
metadata:
author: usestrix
@@ -12,7 +12,7 @@ metadata:
Strix runs autonomous AI pentesting agents that dynamically exploit a target and only report findings validated with a working proof-of-concept. There are **two ways to run it, built on the same engine and producing the same findings** — pick per situation, and mix them freely:
- **Open-source CLI** (self-hosted) — runs on your machine in a Docker sandbox with your own LLM key. Free, fully local, BYO-LLM, air-gap capable. Docs: [docs.strix.ai](https://docs.strix.ai).
- **Cloud API** (managed) — runs on Strix's infrastructure via `https://app.strix.ai/api/v1`. No Docker, no LLM key, no local compute; adds team dashboards, scheduling, PR reviews, downloadable PDF/DOCX reports (Enterprise plan), and internal-network connectors. Docs: [docs.app.strix.ai](https://docs.app.strix.ai). Full workflow in the **strix-cloud-api** skill.
- **Cloud API** (managed) — runs on Strix's infrastructure via `https://app.strix.ai/api/v1`. No Docker, no LLM key, no local compute; adds team dashboards, scheduling, PR reviews, downloadable PDF/DOCX reports (Enterprise plan), and internal-network connectors. Docs: [docs.app.strix.ai](https://docs.app.strix.ai). Full workflow in the **managed-pentesting-with-strix** skill.
## Which one? (decide, don't default)
@@ -112,7 +112,7 @@ Artifacts land in `strix_runs/<run-name>/`:
# Option B — Cloud API (managed, no local infra)
Full details, asset registration, polling, reports, PR reviews, schedules, and webhooks are in the **strix-cloud-api** skill. Minimal launch-and-poll:
Full details, asset registration, polling, reports, PR reviews, schedules, and webhooks are in the **managed-pentesting-with-strix** skill. Minimal launch-and-poll:
```bash
export STRIX_API_TOKEN="<token>" # org-scoped bearer, from Settings → API Access at app.strix.ai
@@ -136,7 +136,7 @@ Ask the user to create the token (and register the target as a domain/repository
## Reporting & next steps
Summarize findings by severity (critical/high/medium/low/info) and include the PoC evidence. To remediate and verify fixes (via either path), use the **strix-fix-findings** skill. To wire scanning into CI/CD, use the **strix-ci-setup** skill.
Summarize findings by severity (critical/high/medium/low/info) and include the PoC evidence. To remediate and verify fixes (via either path), use the **fix-security-vulnerabilities-with-strix** skill. To wire scanning into CI/CD, use the **ci-security-scanning-with-strix** skill.
## Safety
+1 -4
View File
@@ -40,9 +40,6 @@ datas += collect_data_files('tiktoken')
datas += collect_data_files('tiktoken_ext')
datas += collect_data_files('litellm')
# Frozen binaries need certifi's CA bundle on disk; without it TLS to LLM
# providers fails with a generic httpx/litellm "Connection error".
datas += collect_data_files('certifi')
datas += collect_data_files('agents', includes=['**/*.md', '**/*.jinja', '**/*.json'])
@@ -254,7 +251,7 @@ a = Analysis(
hiddenimports=hiddenimports,
hookspath=[],
hooksconfig={},
runtime_hooks=[str(project_root / 'hooks' / 'rthooks' / 'pyi_rth_certifi.py')],
runtime_hooks=[],
excludes=excludes,
noarchive=False,
optimize=0,
+3 -1
View File
@@ -160,7 +160,9 @@ def _schema_types(spec: dict[str, Any]) -> set[str]:
def _decode_structured(value: str, types: set[str]) -> Any:
stripped = value.strip()
if not stripped:
return value
# An empty string is the model's "no value" for a list/dict param; give it
# the empty container so it validates instead of failing the type check.
return [] if "array" in types else {}
try:
decoded = json.loads(stripped)
except json.JSONDecodeError:
+9 -1
View File
@@ -39,6 +39,8 @@ INTERACTIVE BEHAVIOR:
- To end the whole engagement, call the lifecycle tool: finish_scan (root) or agent_finish (subagent).
- A turn that ends with plain text and no tool call does NOT stop you: the system nudges you to continue and will re-run you. Do not rely on going silent to pause — it will not pause you.
- Answering a user question: put the answer in respond_to_user's message. Do not write the answer as plain text and then fall silent — that does not reach a stopping point, it just triggers a continuation nudge.
- If all you want to do is reply and stop, that whole turn is ONE respond_to_user call carrying the answer. Do not write the answer as text and then call respond_to_user as well: the user reads it twice.
- If you do end a turn on plain text and the nudge arrives, your words already reached the user. Do not restate them: call respond_to_user with NO message to simply wait, or with only whatever you still need to add.
- You may include brief explanatory text before a tool call, and you can narrate while you work — plain text is shown to the user as you go. Narrating is free; respond_to_user is specifically the act of WAITING for the user, so do not call it just to give a status update.
- Respond naturally when the user asks questions or gives instructions.
- While actively working on a task, every turn should carry exactly one tool call — use think to plan, the appropriate tool to act, and respond_to_user only when you genuinely need the user.
@@ -261,7 +263,13 @@ Remember: A single well-validated high-impact vulnerability is worth more than d
<multi_agent_system>
AGENT ISOLATION & SANDBOXING:
- All agents run in the same shared Docker container for efficiency
- Each agent has its own: browser sessions, terminal sessions
- Each agent has its own terminal sessions
- Browsers are NOT per-agent by default: `agent-browser` with no `--session` is one
shared browser, so a concurrent agent's navigation invalidates your page and refs.
Pass `--session <your-agent-name>` for any browser work of your own — then it is
yours alone. Each session is a full Chromium (~340 MB) on this shared box, so keep
one, not several, and `agent-browser --session <name> close` when you're done with
the target; an idle browser is reclaimed automatically after 3 minutes
- All agents share the same /workspace directory and proxy history
- Agents can see each other's files and proxy traffic for better collaboration
+19
View File
@@ -733,6 +733,25 @@ def _configure_litellm_default(name: str, value: str) -> None:
setattr(litellm, name, value)
def fallback_model_rejection(fallback: str, primary: str, settings: Settings) -> str | None:
"""Why ``fallback`` cannot stand in for ``primary`` mid-run, or None if it can.
Provider credentials, SDK route, and each agent's tool wrappers are all set
up once from the primary model, so a fallback that needs a different
provider or tool schema would be rejected on every request it serves.
"""
if (
_split_model_provider(_normalized_model_name(fallback))[0]
!= (_split_model_provider(_normalized_model_name(primary))[0])
):
return "needs a different provider, whose credentials are not configured"
if uses_chat_completions_tool_schema(fallback, settings) != uses_chat_completions_tool_schema(
primary, settings
):
return "needs a different tool schema than the agents are built with"
return None
def uses_chat_completions_tool_schema(model_name: str, settings: Settings) -> bool:
"""Return whether the resolved SDK route can only receive JSON function tools."""
if codex.subscription_model(model_name):
+6
View File
@@ -48,6 +48,12 @@ class LlmSettings(BaseSettings):
default=False,
alias="STRIX_FORCE_REQUIRED_TOOL_CHOICE",
)
# A model that serves any content-denied turn (e.g. a ChatGPT-subscription
# cyber-risk guardrail block), and that an agent is pinned to for the rest
# of its lifecycle once it has been denied ``denied_retries`` times. Unset
# disables the fallback: a denial stays terminal for that agent as before.
fallback_model: str | None = Field(default=None, alias="STRIX_LLM_FALLBACK")
denied_retries: int = Field(default=3, ge=0, alias="STRIX_LLM_DENIED_RETRIES")
prompt_cache: bool = Field(
default=True,
alias="STRIX_PROMPT_CACHE",
+54
View File
@@ -18,6 +18,7 @@ if TYPE_CHECKING:
from agents.items import TResponseInputItem
from agents.memory import Session
from agents.model_settings import ModelSettings
logger = logging.getLogger(__name__)
@@ -53,6 +54,11 @@ class AgentCoordinator:
self.errors: dict[str, str] = {}
self.recovery_counts: dict[str, int] = {}
self.idle_resume_counts: dict[str, int] = {}
self.denial_counts: dict[str, int] = {}
self.denial_fallback: set[str] = set()
self._denial_fallback_model: str | None = None
self._denial_fallback_model_settings: ModelSettings | None = None
self._denied_retries: int = 3
self.wait_kinds: dict[str, WaitKind] = {}
self.runtimes: dict[str, AgentRuntime] = {}
self._parent_notified: set[str] = set()
@@ -67,6 +73,29 @@ class AgentCoordinator:
def set_snapshot_path(self, path: Path) -> None:
self._snapshot_path = path
def configure_denial_fallback(
self,
model: str | None,
denied_retries: int,
model_settings: ModelSettings | None = None,
) -> None:
"""Configure the per-agent content denial fallback."""
self._denial_fallback_model = model
self._denial_fallback_model_settings = model_settings
self._denied_retries = denied_retries
@property
def denial_fallback_model(self) -> str | None:
return self._denial_fallback_model
@property
def denial_fallback_model_settings(self) -> ModelSettings | None:
return self._denial_fallback_model_settings
@property
def denied_retries(self) -> int:
return self._denied_retries
def mark_shutting_down(self) -> None:
self.is_shutting_down = True
@@ -243,6 +272,27 @@ class AgentCoordinator:
return
await self._maybe_snapshot()
async def record_denial(self, agent_id: str) -> int:
"""Count a content denial; return the new total."""
async with self._lock:
count = self.denial_counts.get(agent_id, 0) + 1
self.denial_counts[agent_id] = count
await self._maybe_snapshot()
return count
async def mark_denial_fallback(self, agent_id: str) -> None:
"""Mark an agent as using its content denial fallback."""
async with self._lock:
if agent_id in self.denial_fallback:
return
self.denial_fallback.add(agent_id)
await self._maybe_snapshot()
async def is_on_denial_fallback(self, agent_id: str) -> bool:
"""Return whether an agent is using its content denial fallback."""
async with self._lock:
return agent_id in self.denial_fallback
async def set_status(
self, agent_id: str, status: Status | str, *, error: str | None = None
) -> None:
@@ -473,6 +523,8 @@ class AgentCoordinator:
"pending_counts": dict(self.pending_counts),
"recovery_counts": dict(self.recovery_counts),
"idle_resume_counts": dict(self.idle_resume_counts),
"denial_counts": dict(self.denial_counts),
"denial_fallback": sorted(self.denial_fallback),
"wait_kinds": dict(self.wait_kinds),
"mailboxes": {
aid: [dict(m) for m in runtime.mailbox]
@@ -495,6 +547,8 @@ class AgentCoordinator:
self.errors = dict(snap.get("errors", {}))
self.recovery_counts = dict(snap.get("recovery_counts", {}))
self.idle_resume_counts = dict(snap.get("idle_resume_counts", {}))
self.denial_counts = dict(snap.get("denial_counts", {}))
self.denial_fallback = set(snap.get("denial_fallback", []))
self.wait_kinds = dict(snap.get("wait_kinds", {}))
mailboxes = snap.get("mailboxes", {})
if isinstance(mailboxes, dict):
+47 -4
View File
@@ -4,6 +4,7 @@ from __future__ import annotations
import asyncio
import contextlib
import dataclasses
import logging
import uuid
from collections.abc import Callable
@@ -643,7 +644,20 @@ async def _run_cycle( # noqa: PLR0912, PLR0915
image_strips = 0
compactions = 0
model_retries = 0
retry_on_fallback = False
while True:
active_run_config = run_config
if coordinator.denial_fallback_model and (
retry_on_fallback or await coordinator.is_on_denial_fallback(agent_id)
):
active_run_config = dataclasses.replace(
run_config,
model=coordinator.denial_fallback_model,
model_settings=(
coordinator.denial_fallback_model_settings or run_config.model_settings
),
)
retry_on_fallback = False
stream: Any = None
pre_run_items: list[Any] = []
try:
@@ -656,7 +670,7 @@ async def _run_cycle( # noqa: PLR0912, PLR0915
except Exception:
logger.exception("image-budget enforcement failed for %s", agent_id)
try:
await _compact_session(agent, session, run_config, force=False)
await _compact_session(agent, session, active_run_config, force=False)
except Exception:
logger.exception("proactive compaction failed for %s", agent_id)
with contextlib.suppress(Exception):
@@ -664,7 +678,7 @@ async def _run_cycle( # noqa: PLR0912, PLR0915
stream = Runner.run_streamed(
agent,
input=input_data,
run_config=run_config,
run_config=active_run_config,
context=context,
max_turns=max_turns,
session=session,
@@ -744,7 +758,9 @@ async def _run_cycle( # noqa: PLR0912, PLR0915
and is_context_overflow(exc)
):
try:
compacted = await _compact_session(agent, session, run_config, force=True)
compacted = await _compact_session(
agent, session, active_run_config, force=True
)
except Exception:
logger.exception("overflow compaction recovery failed for %s", agent_id)
compacted = False
@@ -757,6 +773,33 @@ async def _run_cycle( # noqa: PLR0912, PLR0915
)
input_data = []
continue
if (
coordinator.denial_fallback_model
and codex.is_content_guardrail_error(exc)
and not await coordinator.is_on_denial_fallback(agent_id)
):
denials = await coordinator.record_denial(agent_id)
retry_on_fallback = True
if denials >= coordinator.denied_retries:
await coordinator.mark_denial_fallback(agent_id)
logger.warning(
"agent %s hit %d content denial(s); pinned to %s for the rest "
"of its lifecycle",
agent_id,
denials,
coordinator.denial_fallback_model,
)
else:
logger.warning(
"agent %s content-denied (%d/%d); replaying this turn on %s",
agent_id,
denials,
coordinator.denied_retries,
coordinator.denial_fallback_model,
)
if session is not None:
input_data = []
continue
if model_retries < _MAX_TRANSIENT_MODEL_RETRIES and _is_transient_model_error(exc):
model_retries += 1
delay = _transient_model_retry_delay(model_retries)
@@ -830,7 +873,7 @@ async def _append_tool_required_message(
"execution and never hands control to the user: it is shown to the user, and the "
"run continues. Continue immediately and call exactly one tool. "
"If you have something to tell the user and nothing to do until they reply, "
"call respond_to_user. "
"call respond_to_user — with no message if you have already said it. "
"If you are blocked waiting for another agent, call wait_for_agents. "
f"If the whole engagement is complete, call {finish_tool}. "
"Otherwise use the appropriate execution or planning tool. "
+11 -1
View File
@@ -138,6 +138,15 @@ def build_root_task(scan_config: dict[str, Any]) -> str:
"target to assess: the instructions below are the only source of "
"truth for what to do."
)
elif not parts and user_instructions:
# Neither a target nor a directory, but there is an instruction: the user
# declined the mount, so the instruction is all there is. Say so, or the
# agent goes looking for a scope that was never given.
parts.append(
"\n\nNo scan target and no working directory were provided. The "
"instructions below are the only source of truth for what to do; "
"work from them and from what you can reach yourself."
)
parts.extend(_render_diff_scope(diff_scope))
@@ -192,9 +201,10 @@ def make_model_settings(
request_timeout: float | None = None,
prompt_cache: bool = True,
extra_headers: dict[str, str] | None = None,
has_tools: bool = True,
) -> ModelSettings:
model_settings = ModelSettings(
parallel_tool_calls=False,
parallel_tool_calls=False if has_tools else None,
retry=DEFAULT_MODEL_RETRY,
include_usage=True,
extra_args=request_timeout_extra_args(request_timeout),
+35 -3
View File
@@ -2,6 +2,7 @@
from __future__ import annotations
import asyncio
import contextlib
import io
import json
@@ -21,6 +22,7 @@ from strix.config import load_settings
from strix.config.models import (
StrixProvider,
configure_sdk_model_defaults,
fallback_model_rejection,
uses_chat_completions_tool_schema,
)
from strix.config.settings import DEFAULT_MAX_TURNS
@@ -174,6 +176,28 @@ async def run_strix_scan(
if coordinator is None:
coordinator = AgentCoordinator()
coordinator.set_snapshot_path(agents_path)
fallback_model = settings.llm.fallback_model
if fallback_model and (
rejection := fallback_model_rejection(fallback_model, resolved_model, settings)
):
raise RuntimeError(
f"STRIX_LLM_FALLBACK '{fallback_model}' {rejection}; it could never serve a "
f"turn for '{resolved_model}'. Pick a fallback from the same provider family."
)
coordinator.configure_denial_fallback(
fallback_model,
settings.llm.denied_retries,
model_settings=make_model_settings(
settings.llm.reasoning_effort,
model_name=fallback_model,
force_required_tool_choice=settings.llm.force_required_tool_choice,
request_timeout=settings.llm.timeout,
prompt_cache=settings.llm.prompt_cache,
extra_headers=settings.llm.extra_headers,
)
if fallback_model
else None,
)
from strix.tools.notes.tools import hydrate_notes_from_disk
from strix.tools.todo.tools import hydrate_todos_from_disk
@@ -429,7 +453,6 @@ async def run_strix_scan(
except BudgetExceededError as exc:
logger.info("Scan %s stopped: %s", scan_id, exc)
if root_id is not None:
await coordinator.cancel_descendants(root_id)
with contextlib.suppress(Exception):
await coordinator.set_status(root_id, "stopped")
return None
@@ -442,19 +465,28 @@ async def run_strix_scan(
scan_id,
)
if root_id is not None:
await coordinator.cancel_descendants(root_id)
with contextlib.suppress(Exception):
await coordinator.set_status(root_id, "stopped")
return None
except (asyncio.CancelledError, KeyboardInterrupt):
logger.info("Scan %s interrupted by the user", scan_id)
if root_id is not None:
with contextlib.suppress(Exception):
await coordinator.set_status(root_id, "running")
raise
except BaseException:
logger.exception("Strix scan %s failed", scan_id)
if root_id is not None:
await coordinator.cancel_descendants(root_id)
with contextlib.suppress(Exception):
await coordinator.set_status(root_id, "failed")
raise
finally:
configure_spill_writer(None)
# Settle descendants before closing sessions: on a clean finish a child
# can still be mid-turn, and closing its session underneath it crashes it.
if root_id is not None:
with contextlib.suppress(Exception):
await coordinator.cancel_descendants(root_id)
for s in sessions_to_close:
with contextlib.suppress(Exception):
s.close()
+20 -2
View File
@@ -4,6 +4,8 @@ from __future__ import annotations
import asyncio
import logging
import sqlite3
from contextlib import contextmanager
from typing import TYPE_CHECKING, Any, cast
from weakref import WeakKeyDictionary
@@ -12,7 +14,7 @@ from agents.memory import SQLiteSession
if TYPE_CHECKING:
from collections.abc import Callable
from collections.abc import Callable, Iterator
from pathlib import Path
from agents.items import TResponseInputItem
@@ -22,9 +24,25 @@ if TYPE_CHECKING:
logger = logging.getLogger(__name__)
class _PooledConnectionSession(SQLiteSession):
@contextmanager
def _locked_connection(self) -> Iterator[sqlite3.Connection]:
with self._lock:
if self._closed:
raise RuntimeError("SQLiteSession is closed")
if self._is_memory_db:
yield self._shared_connection
return
connection = sqlite3.connect(str(self.db_path), check_same_thread=False)
try:
yield connection
finally:
connection.close()
def open_agent_session(agent_id: str, path: Path) -> SQLiteSession:
path.parent.mkdir(parents=True, exist_ok=True)
return SQLiteSession(session_id=agent_id, db_path=path)
return _PooledConnectionSession(session_id=agent_id, db_path=path)
async def seed_initial_input(session: Session, initial_input: Any) -> bool:
+4 -3
View File
@@ -328,10 +328,11 @@ def _load_resume_state(args: argparse.Namespace, parser: argparse.ArgumentParser
parser.error(f"--resume {args.resume}: run.json unreadable: {exc}")
args.targets_info = state.get("targets_info") or []
# A target-less run has no targets_info at all: it works in a mounted
# directory, driven by its instruction.
# A target-less run has no targets_info at all. It is driven by its
# instruction, over a mounted working directory or over nothing when the
# mount was declined, so either of those is enough to resume it.
workspace_mount = state.get("workspace_mount") or None
if not args.targets_info and not workspace_mount:
if not args.targets_info and not workspace_mount and not state.get("user_instruction"):
parser.error(f"--resume {args.resume}: run.json has no targets_info")
for target in args.targets_info:
+7 -37
View File
@@ -6,7 +6,6 @@ Strix Agent Interface
import argparse
import asyncio
import contextlib
import logging
import os
import sys
from pathlib import Path
@@ -43,23 +42,7 @@ from strix.interface.utils import (
build_final_stats_text,
)
from strix.telemetry import posthog, scarf
from strix.telemetry.logging import (
attach_preflight_logging,
configure_dependency_logging,
debug_logging_enabled,
)
# Frozen (PyInstaller) binaries need the bundled certifi CA path exported so
# httpx/requests verify TLS against a real cacert.pem inside the archive.
# The PyInstaller runtime hook covers the official build; this is an extra
# safety net for any frozen entry that loads this module.
if getattr(sys, "frozen", False):
import certifi
_ca_bundle = certifi.where()
os.environ.setdefault("SSL_CERT_FILE", _ca_bundle)
os.environ.setdefault("REQUESTS_CA_BUNDLE", _ca_bundle)
from strix.telemetry.logging import configure_dependency_logging
BEDROCK_MODEL_PREFIX = "bedrock/"
@@ -74,6 +57,9 @@ VERTEX_EXTRA_HINT = (
)
import logging # noqa: E402
logger = logging.getLogger(__name__)
@@ -238,6 +224,7 @@ async def warm_up_llm(show_model_warning: bool = True) -> None:
request_timeout=llm.timeout,
prompt_cache=False,
extra_headers=settings.dedupe.extra_headers,
has_tools=False,
)
if deduper_extra:
merged = {**(deduper_settings.extra_args or {}), **deduper_extra}
@@ -366,30 +353,16 @@ def _print_error_panel(title: str, message: str) -> None:
console.print()
def _format_connection_error_detail(exc: BaseException) -> str:
"""Return the user-facing error detail for a model connection failure.
With ``STRIX_DEBUG`` enabled, include the full ``__cause__`` /
``__context__`` chain so wrapped TLS failures (e.g.
``SSLCertVerificationError`` under litellm/httpx ``Connection error``)
are visible.
"""
if debug_logging_enabled():
return " | ".join(_exception_messages(exc))
return str(exc)
def _print_model_connection_error(exc: BaseException, model_name: str) -> None:
console = Console()
error_text = Text()
detail = _format_connection_error_detail(exc)
sub_hint = _subscription_error_hint(exc)
if sub_hint is not None:
border_style = "yellow"
error_text.append("MODEL NOT AVAILABLE ON SUBSCRIPTION", style="bold yellow")
error_text.append("\n\n", style="white")
error_text.append(f"{sub_hint}\n", style="white")
error_text.append(f"\nDetails: {detail}", style="dim white")
error_text.append(f"\nDetails: {exc}", style="dim white")
else:
border_style = "red"
error_text.append("LLM CONNECTION FAILED", style="bold red")
@@ -399,7 +372,7 @@ def _print_model_connection_error(exc: BaseException, model_name: str) -> None:
hint = _provider_import_hint(exc, model_name)
if hint is not None:
error_text.append(f"\n{hint}\n", style="bold yellow")
error_text.append(f"\nError: {detail}", style="dim white")
error_text.append(f"\nError: {exc}", style="dim white")
panel = Panel(
error_text,
@@ -423,9 +396,6 @@ def _bootstrap_scan(args: argparse.Namespace) -> None:
validate_environment()
if not args.non_interactive:
return
# Preflight runs before prepare_run()/setup_scan_logging(), so attach a
# stderr handler now or STRIX_DEBUG=1 never shows warm-up failures.
attach_preflight_logging()
try:
asyncio.run(warm_up_llm(show_model_warning=True))
except ModelConnectionError as exc:
+1
View File
@@ -78,6 +78,7 @@ async def preflight_model_connection(
request_timeout=resolved_settings.llm.timeout,
prompt_cache=False,
extra_headers=resolved_settings.llm.extra_headers,
has_tools=False,
)
await asyncio.wait_for(
model.get_response(
+5 -14
View File
@@ -138,13 +138,6 @@ class TuiController:
self.error = detail
self.notify_changed()
def enter_setup(self) -> None:
"""Return a session to the start screen, e.g. on a declined mount."""
self.setup_mode = True
self.scan_started = False
self.scan_state = "setup"
self.notify_changed()
def add_message(self, text: str, level: str = "info") -> None:
self._append_message(text, level)
self.notify_changed()
@@ -356,14 +349,12 @@ class TuiController:
if not isinstance(approved, bool):
raise TypeError("approved must be a boolean")
self.pending_workspace_mount = None
if not approved:
# Nothing was prepared, so return to the start screen untouched.
self.workspace_mount = None
self.enter_setup()
return {"approved": False}
self.workspace_mount = mount
# Declining skips the mount, it does not abandon the scan. The prompt is
# the whole of the input either way; the working directory is only an
# extra the agent may look at, so the run goes ahead without one.
self.workspace_mount = mount if approved else None
await self._begin_scan(self._pending_verify)
return {"approved": True}
return {"approved": approved}
async def _send_message(self, payload: dict[str, Any]) -> dict[str, Any]:
agent_id = self._required_string(payload, "agent_id")
+3 -6
View File
@@ -66,13 +66,10 @@ func (m *Model) submitSetupPrompt(value string) (tea.Model, tea.Cmd) {
}
// answerMountConfirmation replies to the working-directory mount the backend is
// waiting on. Declining returns to the start screen, so the prompt goes back in
// the composer to be edited or given a target instead.
// waiting on. Either answer starts the scan - declining only means it runs
// without the directory - so the prompt stays with the run rather than coming
// back to the composer.
func (m *Model) answerMountConfirmation(approved bool) tea.Cmd {
if !approved && m.pendingPrompt != "" {
m.input.SetValue(m.pendingPrompt)
m.resizeViewport()
}
m.pendingPrompt = ""
return send(m.client, "setup.confirm_mount", map[string]any{"approved": approved})
}
@@ -247,13 +247,10 @@ func TestMountConfirmationAnswers(t *testing.T) {
if payload.Approved != tc.approved {
t.Fatalf("%s: approved=%v, want %v", tc.name, payload.Approved, tc.approved)
}
// Declining returns to the start screen, so the prompt comes back.
want := ""
if !tc.approved {
want = "find auth bugs in the login flow"
}
if got := model.input.Value(); got != want {
t.Fatalf("%s: composer = %q, want %q", tc.name, got, want)
// Either answer launches, so the prompt stays with the run rather than
// coming back to the composer.
if got := model.input.Value(); got != "" {
t.Fatalf("%s: composer = %q, want it cleared", tc.name, got)
}
if model.pendingPrompt != "" {
t.Fatalf("%s: held prompt was not cleared: %q", tc.name, model.pendingPrompt)
@@ -290,3 +287,101 @@ func TestSetupPromptWithTargetLaunches(t *testing.T) {
t.Fatalf("setup.start (%d) must come after setup.set_instruction (%d): %v", start, instr, types)
}
}
// The prompt's buttons are buttons: clicking Cancel has to answer the backend,
// which it could not do while the mouse handler had no case for this modal.
func TestMountPromptButtonsAreClickable(t *testing.T) {
for _, testCase := range []struct {
label string
approved bool
}{
{mountConfirmLabel, true},
{mountCancelLabel, false},
} {
connection := &recordingConn{}
model := New(&Client{conn: connection})
model.width, model.height = 130, 40
model.snapshot = protocol.Snapshot{SetupMode: true, WorkingDir: "/Users/me/code/api"}
updated, _ := model.submit("find auth bugs in the login flow")
model = updated.(Model)
connection.Reset()
model.snapshot = protocol.Snapshot{
ScanStarted: true, ScanState: "preparing", PendingMount: "/Users/me/code/api",
}
model.syncMountPrompt()
left, top, panel := model.mountPromptBounds()
clicked := false
for row, line := range strings.Split(panel, "\n") {
plain := ansi.Strip(line)
index := strings.Index(plain, testCase.label)
if index < 0 {
continue
}
updated, cmd := model.updateModalMouse(tea.MouseMsg{
X: left + ansi.StringWidth(plain[:index]) + 1, Y: top + row,
Button: tea.MouseButtonLeft, Action: tea.MouseActionPress,
})
model = updated.(Model)
envelopes := drainCommands(t, cmd, connection)
if len(envelopes) != 1 || envelopes[0].Type != "setup.confirm_mount" {
t.Fatalf("clicking %s sent %v", testCase.label, commandTypes(envelopes))
}
var payload struct {
Approved bool `json:"approved"`
}
if err := json.Unmarshal(envelopes[0].Payload, &payload); err != nil {
t.Fatal(err)
}
if payload.Approved != testCase.approved {
t.Fatalf("clicking %s answered approved=%v", testCase.label, payload.Approved)
}
clicked = true
break
}
if !clicked {
t.Fatalf("%s was not found in the prompt", testCase.label)
}
}
}
// Skipping the mount runs the scan without a directory. It must not throw the
// session back to the start screen, and it must not hand the prompt back: the
// run has it.
func TestSkippingTheMountKeepsTheScanRunning(t *testing.T) {
connection := &recordingConn{}
model := New(&Client{conn: connection})
model.width, model.height = 130, 40
model.snapshot = protocol.Snapshot{SetupMode: true, WorkingDir: "/Users/me/code/api"}
updated, _ := model.submit("find auth bugs in the login flow")
model = updated.(Model)
model.snapshot = protocol.Snapshot{
ScanStarted: true, ScanState: "preparing", PendingMount: "/Users/me/code/api",
}
model.syncMountPrompt()
if model.modal != modalConfirmMount {
t.Fatal("the prompt did not open")
}
model.modalChoice = 1
updated, _ = model.updateModal(tea.KeyMsg{Type: tea.KeyEnter})
model = updated.(Model)
// The backend answers by starting the scan with no mount.
model.handleEnvelope(stateEnvelope(t, 2, protocol.Snapshot{
ScanStarted: true, ScanState: "running",
}))
if model.modal != modalNone {
t.Fatalf("the prompt is still open: %v", model.modal)
}
if model.snapshot.SetupMode {
t.Fatal("skipping the mount fell back to the start screen")
}
if got := model.input.Value(); got != "" {
t.Fatalf("the prompt came back to the composer: %q", got)
}
if model.pendingPrompt != "" {
t.Fatalf("the held prompt was not released: %q", model.pendingPrompt)
}
}
+26 -3
View File
@@ -474,6 +474,18 @@ func (m Model) updateModalMouse(msg tea.MouseMsg) (tea.Model, tea.Cmd) {
m.modalChoice = 1
return m.updateModal(tea.KeyMsg{Type: tea.KeyEnter})
}
case modalConfirmMount:
left, top, panel := m.mountPromptBounds()
if labelHitAt(panel, mountConfirmLabel, left, top, msg.X, msg.Y) {
m.modalChoice = 0
cmd := m.answerMountConfirmation(true)
return m, cmd
}
if labelHitAt(panel, mountCancelLabel, left, top, msg.X, msg.Y) {
m.modalChoice = 1
cmd := m.answerMountConfirmation(false)
return m, cmd
}
case modalVulnerability:
for _, button := range m.reportButtons() {
if button == reportCopy || button == reportDone {
@@ -507,7 +519,14 @@ func (m Model) centeredViewBounds(view string) (left, top, width, height int) {
func (m Model) centeredLabelHit(view, label string, x, y int) bool {
left, top, _, _ := m.centeredViewBounds(view)
for row, line := range strings.Split(view, "\n") {
return labelHitAt(view, label, left, top, x, y)
}
// labelHitAt reports whether a click landed on a label drawn in a panel whose
// top-left corner is at (left, top). The mount prompt is docked in a corner
// rather than centered, so it cannot use the centered bounds.
func labelHitAt(panel, label string, left, top, x, y int) bool {
for row, line := range strings.Split(panel, "\n") {
plain := ansi.Strip(line)
index := strings.Index(plain, label)
if index < 0 || y != top+row {
@@ -593,7 +612,8 @@ func (m Model) updateModal(key tea.KeyMsg) (tea.Model, tea.Cmd) {
case "esc":
if m.modal == modalConfirmMount {
// The backend is waiting on an answer; escape declines it.
return m, m.answerMountConfirmation(false)
cmd := m.answerMountConfirmation(false)
return m, cmd
}
m.closeModal()
return m, nil
@@ -604,7 +624,10 @@ func (m Model) updateModal(key tea.KeyMsg) (tea.Model, tea.Cmd) {
modal, choice := m.modal, m.modalChoice
if modal == modalConfirmMount {
// The snapshot closes this prompt once the backend has the answer.
return m, m.answerMountConfirmation(choice == 0)
// Bound to a variable first: the call restores the held prompt into
// the composer, and that has to be in the model being returned.
cmd := m.answerMountConfirmation(choice == 0)
return m, cmd
}
m.closeModal()
if choice == 1 {
+17
View File
@@ -271,6 +271,23 @@ func (m Model) viewInner() string {
return m.toastOverlay(main)
}
// mountPromptBounds is where the working-directory prompt is drawn. It is placed
// by cornerOverlay rather than centered, so a click has to be tested against
// these bounds and not the ones the other modals use.
func (m Model) mountPromptBounds() (left, top int, panel string) {
panel = m.modalView()
if panel == "" {
return 0, 0, ""
}
_, _, chatWidth, _ := m.layout()
left = max(0, min(chatWidth, m.width)-lipgloss.Width(panel))
statusH := 0
if m.statusVisible() {
statusH = 1
}
return left, max(0, m.inputTop()-statusH-lipgloss.Height(panel)), panel
}
// cornerOverlay splices a panel in directly above the composer, right-aligned
// with it, leaving the rest of the view visible behind it.
func (m Model) cornerOverlay(view, panel string) string {
@@ -220,6 +220,13 @@ func (m Model) confirmView(title string, width int, border, titleColor lipgloss.
return m.confirmDialog(title, "", width, border, titleColor, red, "Yes", "No")
}
// The mount prompt's buttons, named so the renderer and the click test cannot
// drift apart.
const (
mountConfirmLabel = "Mount"
mountCancelLabel = "Skip"
)
// mountConfirmView asks before a target-less scan mounts the working directory.
// It is a compact prompt docked in the corner of the live view: nothing is
// prepared until it is answered, and the directory is a workspace rather than a
@@ -232,8 +239,8 @@ func (m Model) mountConfirmView() string {
}
title := render.Bold(amber).Render("△ Mount working directory?")
body := render.Col(white).Render(truncatePath(dir, width-4)) + "\n" +
render.Dim().Render("writable in the sandbox")
return m.cornerPrompt(title, body, width, "Confirm", "Cancel")
render.Dim().Render("writable in the sandbox · skip to run without it")
return m.cornerPrompt(title, body, width, mountConfirmLabel, mountCancelLabel)
}
// truncatePath keeps the tail of a path visible, which is the part that
+1
View File
@@ -294,6 +294,7 @@ async def _summarize(model: str, prompt: str, max_tokens: int) -> str | None:
request_timeout=llm.timeout,
prompt_cache=False,
extra_headers=llm.extra_headers,
has_tools=False,
).resolve(ModelSettings(max_tokens=max_tokens))
try:
response = (
+1
View File
@@ -62,6 +62,7 @@ def _dedupe_model_settings(
# must never receive the main endpoint's credentials. A dedicated model
# gets its own DEDUPE_LLM_EXTRA_HEADERS instead.
extra_headers=dedupe.extra_headers if dedupe.model else llm.extra_headers,
has_tools=False,
)
extra = _dedupe_extra_args(dedupe)
if extra:
+35 -2
View File
@@ -58,6 +58,26 @@ agent-browser screenshot
The browser stays running across commands so these feel like a single
session. Use `agent-browser close` (or `close --all`) when you're done.
The default session is **shared with every other agent in the sandbox** — if
another agent navigates it, your page and your refs are gone from under you. So
claim your own by passing `--session <your-agent-name>` on **every** command:
```bash
agent-browser --session recon-3 open https://example.com
agent-browser --session recon-3 snapshot -i
agent-browser --session recon-3 close # when done with the target
```
The examples in the rest of this skill omit `--session` to keep them readable;
keep passing yours. Each session is a separate Chromium (~340 MB) on a shared
box, so hold one rather than several, and close it when you're finished.
A browser left idle for 3 minutes is reclaimed automatically to free memory for
the other agents; the next command relaunches it, but the page, tabs, refs and
cookies are gone. If you're authenticated and about to go do something else for a
while, save the state first (see
[Persist session across runs](#persist-session-across-runs)).
## Reading a page
```bash
@@ -307,6 +327,16 @@ agent-browser --session b fill @e1 "bob@test.com"
`AGENT_BROWSER_SESSION=myapp` sets the default session for the current
shell.
Use a session named after yourself for your own work — that's what keeps a
concurrent agent from navigating the page out from under you. Every session is a
separate Chromium though, so hold one at a time rather than a collection, and
close each one when its flow is finished:
```bash
agent-browser --session a close
agent-browser --session b close
```
### Mock network requests
```bash
@@ -368,8 +398,11 @@ agent-browser dialog dismiss # cancel
## Readiness & recovery
The first `agent-browser open` in a session launches the headless-Chrome
daemon; later commands reuse it. Distinguish the two failure modes and react
differently — do **not** blindly re-run the same failing command in a loop:
daemon; later commands reuse it. A daemon left idle for 3 minutes shuts itself
down to free memory for the other agents, so an `open` after a long gap is a
fresh browser rather than a resumed one — expect to re-navigate, and re-`state
load` if you were logged in. Distinguish the failure modes and react differently
— do **not** blindly re-run the same failing command in a loop:
- **Daemon / connection failure** (`Failed to connect`, `connection refused`,
socket missing, `browser not running`): the daemon isn't up or has died. Run
+7 -65
View File
@@ -116,69 +116,6 @@ def _silence_urllib3_finalizer_noise() -> None:
sys.unraisablehook = hook
_DEBUG_ENV_TRUTHY = frozenset({"1", "true", "yes", "on"})
_PREFLIGHT_HANDLER_TAG = "_strix_preflight_handler"
def debug_logging_enabled(*, debug: bool | None = None) -> bool:
"""Resolve whether Strix debug logging is on.
``None`` (default) reads ``STRIX_DEBUG``: ``1`` / ``true`` / ``yes`` /
``on`` (case-insensitive) enables debug.
"""
if debug is not None:
return debug
return (os.environ.get("STRIX_DEBUG") or "").strip().lower() in _DEBUG_ENV_TRUTHY
def attach_preflight_logging(*, debug: bool | None = None) -> None:
"""Attach a stderr-only handler so LLM preflight logs are visible early.
``warm_up_llm`` runs before ``setup_scan_logging`` (which needs a run
directory). Without this, ``STRIX_DEBUG=1`` still produces no output for
preflight failures.
"""
configure_dependency_logging()
enabled = debug_logging_enabled(debug=debug)
level = logging.DEBUG if enabled else logging.ERROR
formatter = logging.Formatter(_FORMAT, datefmt=_DATEFMT)
context_filter = _StrixContextFilter()
stream_handler = logging.StreamHandler()
stream_handler.setLevel(level)
stream_handler.setFormatter(formatter)
stream_handler.addFilter(context_filter)
stream_handler.addFilter(_StdoutQuietFilter())
setattr(stream_handler, _PREFLIGHT_HANDLER_TAG, True)
for name in _TRACKED_ROOTS:
tracked = logging.getLogger(name)
# Replace a previous preflight handler so repeated calls stay idempotent.
for handler in list(tracked.handlers):
if getattr(handler, _PREFLIGHT_HANDLER_TAG, False):
tracked.removeHandler(handler)
with contextlib.suppress(Exception):
handler.close()
tracked.setLevel(logging.DEBUG)
tracked.addHandler(stream_handler)
tracked.propagate = False
for name in _NOISY_LIBS:
logging.getLogger(name).setLevel(logging.WARNING)
def remove_preflight_logging() -> None:
"""Detach any preflight stderr handler from the tracked logger roots."""
for name in _TRACKED_ROOTS:
tracked = logging.getLogger(name)
for handler in list(tracked.handlers):
if getattr(handler, _PREFLIGHT_HANDLER_TAG, False):
tracked.removeHandler(handler)
with contextlib.suppress(Exception):
handler.close()
def setup_scan_logging(run_dir: Path, *, debug: bool | None = None) -> Callable[[], None]:
"""Attach scan-scoped handlers; return a teardown callable.
@@ -196,9 +133,14 @@ def setup_scan_logging(run_dir: Path, *, debug: bool | None = None) -> Callable[
time. Safe to call from a ``finally`` block.
"""
configure_dependency_logging()
remove_preflight_logging()
debug = debug_logging_enabled(debug=debug)
if debug is None:
debug = (os.environ.get("STRIX_DEBUG") or "").strip().lower() in {
"1",
"true",
"yes",
"on",
}
run_dir.mkdir(parents=True, exist_ok=True)
log_path = run_dir / "strix.log"
+5 -1
View File
@@ -15,7 +15,7 @@ def _ctx(ctx: RunContextWrapper) -> dict[str, Any]:
@function_tool
async def respond_to_user(ctx: RunContextWrapper, message: str) -> str:
async def respond_to_user(ctx: RunContextWrapper, message: str = "") -> str:
"""Answer the user and hand control back to them.
This is the ONLY way to yield to the user. Delivering the message and
@@ -45,6 +45,10 @@ async def respond_to_user(ctx: RunContextWrapper, message: str) -> str:
have followed the tool calls that led here. Lead with the
answer or the decision you need, and if you are blocked, say
exactly what you need from them.
Omit it when you have just said your piece as plain text and
only need to wait: that text has already reached them, and
repeating it makes them read the same answer twice.
"""
inner = _ctx(ctx)
coordinator = coordinator_from_context(inner)
+27 -6
View File
@@ -110,12 +110,19 @@ def _get_agent_todos(agent_id: str) -> dict[str, dict[str, Any]]:
def _normalize_priority(priority: str | None, default: str = "normal") -> str:
candidate = (priority or default or "normal").lower()
candidate = str(priority or default or "normal").strip().lower()
if candidate not in VALID_PRIORITIES:
raise ValueError(f"Invalid priority. Must be one of: {', '.join(VALID_PRIORITIES)}")
return candidate
def _coerce_priority(priority: str | None, default: str = "normal") -> str:
try:
return _normalize_priority(priority, default)
except ValueError:
return default
def _sorted_todos(agent_id: str) -> list[dict[str, Any]]:
todos_list = [
{**todo, "todo_id": todo_id} for todo_id, todo in _get_agent_todos(agent_id).items()
@@ -285,11 +292,16 @@ async def create_todo(ctx: RunContextWrapper, todos: str) -> str:
- ``description`` (str, optional): extra context or
acceptance criteria.
- ``priority`` (str, optional): one of ``"low"`` /
``"normal"`` / ``"high"`` / ``"critical"``. Defaults to
``"normal"``.
``"normal"`` / ``"high"`` / ``"critical"``. Anything else,
including omitting it, falls back to ``"normal"`` rather
than failing.
Example: ``[{"title": "Probe /admin", "priority": "high"},
{"title": "Check JWT alg=none"}]``.
A title already on the list, or repeated within this call, is
skipped rather than duplicated; skipped titles come back under
``skipped``.
"""
agent_id = _agent_id_from(ctx)
try:
@@ -302,13 +314,21 @@ async def create_todo(ctx: RunContextWrapper, todos: str) -> str:
)
agent_todos = _get_agent_todos(agent_id)
seen = {todo["title"].strip().lower() for todo in agent_todos.values()}
created: list[dict[str, Any]] = []
skipped: list[dict[str, str]] = []
for task in tasks:
task_priority = _normalize_priority(task.get("priority"))
title = task["title"]
key = title.lower()
if key in seen:
skipped.append({"title": title, "reason": "duplicate title"})
continue
seen.add(key)
task_priority = _coerce_priority(task.get("priority"))
todo_id = str(uuid.uuid4())[:6]
timestamp = datetime.now(UTC).isoformat()
agent_todos[todo_id] = {
"title": task["title"],
"title": title,
"description": task.get("description"),
"priority": task_priority,
"status": "pending",
@@ -316,7 +336,7 @@ async def create_todo(ctx: RunContextWrapper, todos: str) -> str:
"updated_at": timestamp,
"completed_at": None,
}
created.append({"todo_id": todo_id, "title": task["title"], "priority": task_priority})
created.append({"todo_id": todo_id, "title": title, "priority": task_priority})
except (ValueError, TypeError) as e:
return json.dumps(
{"success": False, "error": f"Failed to create todo: {e}"},
@@ -330,6 +350,7 @@ async def create_todo(ctx: RunContextWrapper, todos: str) -> str:
"success": True,
"created": created,
"created_count": len(created),
"skipped": skipped,
"todos": _sorted_todos(agent_id),
"total_count": len(_get_agent_todos(agent_id)),
},
+23 -1
View File
@@ -70,7 +70,6 @@ async def test_encoded_list_is_decoded_for_an_array_parameter(schema: dict[str,
"auth",
"Endpoint /admin leaks user data, and session tokens never expire",
'"auth"',
"",
],
)
async def test_free_form_strings_are_never_split_into_an_array(value: str) -> None:
@@ -79,6 +78,29 @@ async def test_free_form_strings_are_never_split_into_an_array(value: str) -> No
assert parsed["tags"] == value
@pytest.mark.asyncio
@pytest.mark.parametrize("schema", [_ARRAY, _NULLABLE_ARRAY])
@pytest.mark.parametrize("value", ["", " "])
async def test_empty_string_becomes_an_empty_array(schema: dict[str, Any], value: str) -> None:
parsed = await _roundtrip(schema, {"tags": value})
assert parsed["tags"] == []
@pytest.mark.asyncio
async def test_empty_string_becomes_an_empty_object() -> None:
parsed = await _roundtrip(_OBJECT, {"modifications": ""})
assert parsed["modifications"] == {}
@pytest.mark.asyncio
async def test_empty_string_for_a_string_parameter_is_untouched() -> None:
parsed = await _roundtrip(_STRING, {"todos": ""})
assert parsed["todos"] == ""
@pytest.mark.asyncio
async def test_encoded_mapping_is_decoded_for_an_object_parameter() -> None:
parsed = await _roundtrip(_OBJECT, {"modifications": '{"method": "POST"}'})
+168
View File
@@ -0,0 +1,168 @@
from __future__ import annotations
from typing import Any
import pytest
from agents import ModelSettings, RunConfig, Runner
from strix.config import codex
from strix.core import execution
from strix.core.agents import AgentCoordinator
class _FakeStream:
def __init__(self, exc: BaseException | None = None) -> None:
self._exc = exc
self.run_loop_exception: BaseException | None = None
async def stream_events(self) -> Any:
if self._exc is not None:
raise self._exc
events: list[Any] = []
for event in events:
yield event
def _patch_fast_backoff(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(execution, "_TRANSIENT_MODEL_RETRY_BASE_DELAY_S", 0.0)
monkeypatch.setattr(execution, "_TRANSIENT_MODEL_RETRY_MAX_DELAY_S", 0.0)
def _guardrail_stream() -> _FakeStream:
return _FakeStream(codex.CodexContentGuardrailError("gpt-5.6-sol"))
async def _run_once(
monkeypatch: pytest.MonkeyPatch,
streams: list[_FakeStream],
*,
fallback_model: str | None = None,
denied_retries: int = 3,
primary_model: str = "openai/gpt-5.6-sol",
fallback_model_settings: ModelSettings | None = None,
) -> tuple[Any, list[tuple[str | None, ModelSettings]], AgentCoordinator]:
_patch_fast_backoff(monkeypatch)
calls: list[tuple[str | None, ModelSettings]] = []
def _fake_run_streamed(*_args: Any, **kwargs: Any) -> _FakeStream:
run_config = kwargs["run_config"]
calls.append((run_config.model, run_config.model_settings))
return streams[len(calls) - 1]
monkeypatch.setattr(Runner, "run_streamed", _fake_run_streamed)
coordinator = AgentCoordinator()
await coordinator.register("root", "strix", parent_id=None)
if fallback_model is not None:
coordinator.configure_denial_fallback(
fallback_model, denied_retries, model_settings=fallback_model_settings
)
result = await execution._run_cycle(
object(),
coordinator,
"root",
input_data="task",
run_config=RunConfig(model=primary_model, model_settings=ModelSettings()),
context={},
max_turns=5,
session=None,
interactive=False,
event_sink=None,
hooks=None,
)
return result, calls, coordinator
@pytest.mark.asyncio
async def test_a_denied_turn_is_retried_on_the_fallback_model(
monkeypatch: pytest.MonkeyPatch,
) -> None:
streams = [_guardrail_stream(), _FakeStream()]
result, models, coordinator = await _run_once(
monkeypatch,
streams,
fallback_model="openai/gpt-5.4",
)
assert result is streams[1]
assert [model for model, _ in models] == ["openai/gpt-5.6-sol", "openai/gpt-5.4"]
# One denial is below the threshold, so the agent is not pinned to the
# fallback and its next turn starts on the main model again.
assert await coordinator.is_on_denial_fallback("root") is False
@pytest.mark.asyncio
async def test_agent_is_pinned_to_the_fallback_after_repeated_denials(
monkeypatch: pytest.MonkeyPatch,
) -> None:
streams = [_guardrail_stream() for _ in range(3)] + [_FakeStream()]
result, models, coordinator = await _run_once(
monkeypatch,
streams,
fallback_model="openai/gpt-5.4",
)
assert result is streams[3]
assert [model for model, _ in models] == ["openai/gpt-5.6-sol"] + ["openai/gpt-5.4"] * 3
assert await coordinator.is_on_denial_fallback("root") is True
@pytest.mark.asyncio
async def test_run_cycle_does_not_retry_guardrail_without_fallback(
monkeypatch: pytest.MonkeyPatch,
) -> None:
guardrail = codex.CodexContentGuardrailError("gpt-5.6-sol")
with pytest.raises(codex.CodexContentGuardrailError):
await _run_once(monkeypatch, [_FakeStream(guardrail), _FakeStream()])
@pytest.mark.asyncio
async def test_run_cycle_pins_on_first_denial_at_boundary(
monkeypatch: pytest.MonkeyPatch,
) -> None:
streams = [_guardrail_stream(), _FakeStream()]
result, models, coordinator = await _run_once(
monkeypatch,
streams,
fallback_model="openai/gpt-5.4",
denied_retries=1,
)
assert result is streams[1]
assert [model for model, _ in models] == ["openai/gpt-5.6-sol", "openai/gpt-5.4"]
assert await coordinator.is_on_denial_fallback("root") is True
@pytest.mark.asyncio
async def test_fallback_uses_its_own_model_settings(
monkeypatch: pytest.MonkeyPatch,
) -> None:
fallback_settings = ModelSettings(parallel_tool_calls=True)
streams = [_guardrail_stream(), _FakeStream()]
result, calls, _coordinator = await _run_once(
monkeypatch,
streams,
fallback_model="openai/gpt-5.4",
denied_retries=1,
fallback_model_settings=fallback_settings,
)
assert result is streams[1]
assert calls[0][1] is not fallback_settings
assert calls[1][1] is fallback_settings
@pytest.mark.asyncio
async def test_denial_fallback_state_round_trips_through_snapshot() -> None:
coordinator = AgentCoordinator()
await coordinator.register("root", "strix", parent_id=None)
coordinator.configure_denial_fallback("openai/gpt-5.4", 3)
await coordinator.record_denial("root")
await coordinator.mark_denial_fallback("root")
restored = AgentCoordinator()
await restored.restore(await coordinator.snapshot())
assert restored.denial_counts == {"root": 1}
assert await restored.is_on_denial_fallback("root") is True
+37
View File
@@ -1228,3 +1228,40 @@ async def test_wait_kind_survives_a_snapshot_round_trip() -> None:
assert restored.wait_kinds["root"] == "user"
assert restored.idle_resume_counts["root"] == 1
assert await execution._plain_waiting_timeout(restored, "root") is None
@pytest.mark.asyncio
async def test_interactive_nudge_offers_waiting_without_repeating() -> None:
"""The nudge is the instruction an agent reads when it is stranded here.
It is where the option to wait on what was already said has to be, not only
in the system prompt: an agent that ended a turn on plain text reasons off
this text, and without the clause it restates its answer to reach a tool
call, so the user reads it twice.
The clause holds whatever the turn did, because the agent is the one who
knows whether it spoke this fires for a turn that produced no text at all.
"""
items = await execution._append_tool_required_message(
session=None,
context={"parent_id": None},
attempt=1,
limit=3,
interactive=True,
)
assert "with no message if you have already said it" in items[0]["content"]
@pytest.mark.asyncio
async def test_autonomous_nudge_does_not_offer_the_user() -> None:
"""There is nobody attached to an autonomous run to wait for."""
items = await execution._append_tool_required_message(
session=None,
context={"parent_id": None},
attempt=1,
limit=3,
interactive=False,
)
assert "respond_to_user" not in items[0]["content"]
+10
View File
@@ -299,6 +299,16 @@ def test_make_model_settings_forces_required_for_anyllm_routed_openai_model() ->
assert settings.tool_choice == "required"
def test_make_model_settings_disables_parallel_tool_calls_by_default() -> None:
assert make_model_settings("none", model_name="gpt-4o").parallel_tool_calls is False
def test_make_model_settings_omits_parallel_tool_calls_without_tools() -> None:
settings = make_model_settings("none", model_name="gpt-4o", has_tools=False)
assert settings.parallel_tool_calls is None
def test_make_model_settings_sets_request_timeout() -> None:
settings = make_model_settings(
"none",
-14
View File
@@ -9,8 +9,6 @@ import pytest
PROJECT_ROOT = Path(__file__).resolve().parents[1]
SPEC_PATH = PROJECT_ROOT / "strix.spec"
CERTIFI_RTHOOK = PROJECT_ROOT / "hooks" / "rthooks" / "pyi_rth_certifi.py"
def test_wheel_build_requires_go(tmp_path: Path) -> None:
@@ -31,15 +29,3 @@ def test_wheel_build_requires_go(tmp_path: Path) -> None:
assert result.returncode != 0
assert "Go 1.24 or newer is required" in result.stdout + result.stderr
def test_pyinstaller_spec_bundles_certifi_ca_and_runtime_hook() -> None:
spec = SPEC_PATH.read_text(encoding="utf-8")
assert "collect_data_files('certifi')" in spec
assert "pyi_rth_certifi.py" in spec
assert "runtime_hooks=[" in spec
assert CERTIFI_RTHOOK.is_file()
hook = CERTIFI_RTHOOK.read_text(encoding="utf-8")
assert "SSL_CERT_FILE" in hook
assert "REQUESTS_CA_BUNDLE" in hook
assert "certifi.where()" in hook
-102
View File
@@ -1,102 +0,0 @@
"""Tests for exception-chain helpers and preflight debug logging."""
from __future__ import annotations
import logging
import ssl
from typing import TYPE_CHECKING
from strix.interface.main import (
_exception_messages,
_format_connection_error_detail,
)
from strix.telemetry import logging as tlog
from strix.telemetry.logging import (
attach_preflight_logging,
debug_logging_enabled,
remove_preflight_logging,
setup_scan_logging,
)
if TYPE_CHECKING:
from pathlib import Path
import pytest
def _preflight_handlers(name: str) -> list[logging.Handler]:
return [
handler
for handler in logging.getLogger(name).handlers
if getattr(handler, tlog._PREFLIGHT_HANDLER_TAG, False)
]
def test_exception_messages_walks_cause_chain_to_ssl_error() -> None:
root = ssl.SSLCertVerificationError("certificate verify failed")
middle = ConnectionError("TLS handshake failed")
middle.__cause__ = root
exc = ConnectionError("Connection error.")
exc.__cause__ = middle
messages = _exception_messages(exc)
assert "Connection error." in messages
assert "TLS handshake failed" in messages
assert any("certificate verify failed" in message for message in messages)
def test_format_connection_error_detail_includes_chain_when_debug(
monkeypatch: pytest.MonkeyPatch,
) -> None:
root = ssl.SSLCertVerificationError("certificate verify failed")
exc = ConnectionError("Connection error.")
exc.__cause__ = root
monkeypatch.delenv("STRIX_DEBUG", raising=False)
assert _format_connection_error_detail(exc) == "Connection error."
monkeypatch.setenv("STRIX_DEBUG", "1")
detail = _format_connection_error_detail(exc)
assert "Connection error." in detail
assert "certificate verify failed" in detail
def test_debug_logging_enabled_reads_strix_debug(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.delenv("STRIX_DEBUG", raising=False)
assert debug_logging_enabled() is False
assert debug_logging_enabled(debug=True) is True
monkeypatch.setenv("STRIX_DEBUG", "yes")
assert debug_logging_enabled() is True
def test_attach_preflight_logging_emits_debug_to_stderr(
monkeypatch: pytest.MonkeyPatch,
capsys: pytest.CaptureFixture[str],
) -> None:
monkeypatch.setenv("STRIX_DEBUG", "1")
try:
attach_preflight_logging()
logging.getLogger("strix").debug("LLM warm-up failed")
captured = capsys.readouterr()
assert "LLM warm-up failed" in captured.err
finally:
remove_preflight_logging()
def test_setup_scan_logging_removes_preflight_handler(
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
) -> None:
monkeypatch.delenv("STRIX_DEBUG", raising=False)
attach_preflight_logging()
assert _preflight_handlers("strix")
teardown = setup_scan_logging(tmp_path)
try:
for name in ("strix", "openai.agents"):
assert not _preflight_handlers(name)
finally:
teardown()
+28
View File
@@ -64,3 +64,31 @@ async def test_a_message_that_already_arrived_is_taken_instead_of_parking() -> N
assert result["wait_outcome"] == "message_arrived"
assert result["pending_messages"] == 1
assert coordinator.statuses["root"] == "running"
async def _call_without_message(context: dict[str, Any]) -> dict[str, Any]:
ctx = ToolContext(
context=context,
tool_name="respond_to_user",
tool_call_id="call-1",
tool_arguments="{}",
)
raw = await respond_to_user.on_invoke_tool(ctx, "{}")
return json.loads(raw) # type: ignore[no-any-return]
@pytest.mark.asyncio
async def test_parks_without_a_message() -> None:
"""An agent that has already said its piece as plain text can just wait.
The nudge is what leaves it here, and while a message was required the only
way to stop was to send the same answer a second time.
"""
context = await _context(interactive=True)
result = await _call_without_message(context)
assert result["success"] is True
assert result["wait_outcome"] == "waiting"
assert result["message"] == ""
assert context["coordinator"].statuses["root"] == "waiting"
+110
View File
@@ -0,0 +1,110 @@
from __future__ import annotations
import asyncio
import types
from typing import Any
import pytest
from agents import ModelSettings
import strix.tools.notes.tools as notes_tools
import strix.tools.todo.tools as todo_tools
from strix.core import runner
from strix.core.agents import AgentCoordinator
from strix.runtime import session_manager
def _wire_runner(monkeypatch: pytest.MonkeyPatch, tmp_path: Any) -> None:
monkeypatch.setattr(runner, "run_dir_for", lambda _scan_id: tmp_path)
monkeypatch.setattr(runner, "runtime_state_dir", lambda _run_dir: tmp_path)
monkeypatch.setattr(runner, "setup_scan_logging", lambda _run_dir: lambda: None)
monkeypatch.setattr(runner, "set_scan_id", lambda _scan_id: None)
settings = types.SimpleNamespace(
llm=types.SimpleNamespace(
model="openai/gpt-4o",
reasoning_effort="high",
force_required_tool_choice=False,
timeout=300,
prompt_cache=True,
extra_headers=None,
fallback_model=None,
denied_retries=3,
),
runtime=types.SimpleNamespace(max_context_images=3),
)
monkeypatch.setattr(runner, "load_settings", lambda: settings)
monkeypatch.setattr(runner, "configure_sdk_model_defaults", lambda _settings: None)
monkeypatch.setattr(
runner, "uses_chat_completions_tool_schema", lambda _model, _settings: False
)
monkeypatch.setattr(todo_tools, "hydrate_todos_from_disk", lambda _state_dir: None)
monkeypatch.setattr(notes_tools, "hydrate_notes_from_disk", lambda _state_dir: None)
async def _create_or_reuse(*_args: Any, **_kwargs: Any) -> dict[str, Any]:
return {"client": object(), "session": object(), "caido_client": None}
async def _cleanup(*_args: Any, **_kwargs: Any) -> None:
return None
monkeypatch.setattr(session_manager, "create_or_reuse", _create_or_reuse)
monkeypatch.setattr(session_manager, "cleanup", _cleanup)
monkeypatch.setattr(runner, "build_root_task", lambda _scan_config: "task")
monkeypatch.setattr(runner, "build_scope_context", lambda _scan_config: "")
monkeypatch.setattr(runner, "make_model_settings", lambda *_a, **_k: ModelSettings())
monkeypatch.setattr(runner, "build_strix_agent", lambda **_kwargs: object())
monkeypatch.setattr(runner, "make_child_factory", lambda **_kwargs: lambda **_k: object())
monkeypatch.setattr(runner, "open_agent_session", lambda _root_id, _db: object())
def _root_status(coordinator: AgentCoordinator) -> str:
roots = [aid for aid, parent in coordinator.parent_of.items() if parent is None]
assert len(roots) == 1
return coordinator.statuses[roots[0]]
@pytest.mark.parametrize("interrupt", [KeyboardInterrupt, asyncio.CancelledError])
@pytest.mark.asyncio
async def test_user_interrupt_leaves_the_root_running_for_resume(
monkeypatch: pytest.MonkeyPatch, tmp_path: Any, interrupt: type[BaseException]
) -> None:
_wire_runner(monkeypatch, tmp_path)
async def _interrupt(*_args: Any, **_kwargs: Any) -> None:
raise interrupt()
monkeypatch.setattr(runner, "run_agent_loop", _interrupt)
coordinator = AgentCoordinator()
with pytest.raises(interrupt):
await runner.run_strix_scan(
scan_config={"targets": [], "scan_mode": "deep"},
scan_id="scan-test",
image="img",
coordinator=coordinator,
)
assert _root_status(coordinator) == "running"
@pytest.mark.asyncio
async def test_a_real_crash_still_marks_root_failed(
monkeypatch: pytest.MonkeyPatch, tmp_path: Any
) -> None:
_wire_runner(monkeypatch, tmp_path)
async def _boom(*_args: Any, **_kwargs: Any) -> None:
raise RuntimeError("boom")
monkeypatch.setattr(runner, "run_agent_loop", _boom)
coordinator = AgentCoordinator()
with pytest.raises(RuntimeError, match="boom"):
await runner.run_strix_scan(
scan_config={"targets": [], "scan_mode": "deep"},
scan_id="scan-test",
image="img",
coordinator=coordinator,
)
assert _root_status(coordinator) == "failed"
+2
View File
@@ -42,6 +42,8 @@ async def test_persistent_rate_limit_stops_gracefully(
timeout=300,
prompt_cache=True,
extra_headers=None,
fallback_model=None,
denied_retries=3,
),
runtime=types.SimpleNamespace(max_context_images=3),
)
+2
View File
@@ -50,6 +50,8 @@ def _patch_engine_scaffold(
timeout=300,
prompt_cache=True,
extra_headers=None,
fallback_model=None,
denied_retries=3,
),
runtime=types.SimpleNamespace(max_context_images=3),
)
+95
View File
@@ -0,0 +1,95 @@
from __future__ import annotations
import asyncio
import types
from typing import Any
import pytest
from agents import ModelSettings
import strix.tools.notes.tools as notes_tools
import strix.tools.todo.tools as todo_tools
from strix.core import runner
from strix.core.agents import AgentCoordinator
from strix.runtime import session_manager
def _wire_runner(monkeypatch: pytest.MonkeyPatch, tmp_path: Any) -> None:
monkeypatch.setattr(runner, "run_dir_for", lambda _scan_id: tmp_path)
monkeypatch.setattr(runner, "runtime_state_dir", lambda _run_dir: tmp_path)
monkeypatch.setattr(runner, "setup_scan_logging", lambda _run_dir: lambda: None)
monkeypatch.setattr(runner, "set_scan_id", lambda _scan_id: None)
settings = _settings()
monkeypatch.setattr(runner, "load_settings", lambda: settings)
monkeypatch.setattr(runner, "configure_sdk_model_defaults", lambda _s: None)
monkeypatch.setattr(runner, "uses_chat_completions_tool_schema", lambda _m, _s: False)
monkeypatch.setattr(todo_tools, "hydrate_todos_from_disk", lambda _d: None)
monkeypatch.setattr(notes_tools, "hydrate_notes_from_disk", lambda _d: None)
async def _create_or_reuse(*_a: Any, **_k: Any) -> dict[str, Any]:
return {"client": object(), "session": object(), "caido_client": None}
async def _cleanup(*_a: Any, **_k: Any) -> None:
return None
monkeypatch.setattr(session_manager, "create_or_reuse", _create_or_reuse)
monkeypatch.setattr(session_manager, "cleanup", _cleanup)
monkeypatch.setattr(runner, "build_root_task", lambda _c: "task")
monkeypatch.setattr(runner, "build_scope_context", lambda _c: "")
monkeypatch.setattr(runner, "make_model_settings", lambda *_a, **_k: ModelSettings())
monkeypatch.setattr(runner, "build_strix_agent", lambda **_k: object())
monkeypatch.setattr(runner, "make_child_factory", lambda **_k: lambda **_kk: object())
monkeypatch.setattr(runner, "open_agent_session", lambda _root_id, _db: object())
def _settings() -> Any:
return types.SimpleNamespace(
llm=types.SimpleNamespace(
model="openai/gpt-4o",
reasoning_effort="high",
force_required_tool_choice=False,
timeout=300,
prompt_cache=True,
extra_headers=None,
fallback_model=None,
denied_retries=3,
),
runtime=types.SimpleNamespace(max_context_images=3),
)
@pytest.mark.asyncio
async def test_a_live_child_is_settled_before_sessions_close(
monkeypatch: pytest.MonkeyPatch, tmp_path: Any
) -> None:
_wire_runner(monkeypatch, tmp_path)
coordinator = AgentCoordinator()
child_started = asyncio.Event()
child_task: dict[str, asyncio.Task[None]] = {}
async def _root_finishes(**kwargs: Any) -> None:
root_id = kwargs["agent_id"]
async def _child_mid_turn() -> None:
child_started.set()
await asyncio.sleep(3600)
await coordinator.register("child", "Child", parent_id=root_id)
task = asyncio.create_task(_child_mid_turn())
child_task["t"] = task
await coordinator.attach_runtime("child", task=task)
await child_started.wait()
monkeypatch.setattr(runner, "run_agent_loop", _root_finishes)
await runner.run_strix_scan(
scan_config={"targets": [], "scan_mode": "deep"},
scan_id="scan-test",
image="img",
coordinator=coordinator,
)
task = child_task["t"]
assert task.done(), "the child task was left running past scan teardown"
assert task.cancelled(), "the child was not cancelled cleanly on a finish"
+117
View File
@@ -0,0 +1,117 @@
from __future__ import annotations
import asyncio
from pathlib import Path
from typing import Any, cast
import pytest
from strix.core.sessions import open_agent_session
def _count_open_fds() -> int | None:
for path in (Path("/proc/self/fd"), Path("/dev/fd")):
if path.is_dir():
return len(list(path.iterdir()))
return None
@pytest.mark.asyncio
async def test_sessions_hold_no_descriptors_while_parked(tmp_path: Path) -> None:
"""Descriptor use must track live operations, not the number of sessions.
The SDK keeps a connection per (session, pool thread) open for the session's
whole life. An agent parks rather than exits, so its session lives for the
scan, and fan-out multiplies those handles until the process runs out of file
descriptors (#1018). A session that is not mid-operation should hold none.
"""
baseline = _count_open_fds()
if baseline is None:
pytest.skip("no /proc/self/fd or /dev/fd on this platform")
sessions = [open_agent_session(f"a{i}", tmp_path / f"s{i}.db") for i in range(60)]
try:
for _ in range(4):
await asyncio.gather(
*(s.add_items([{"role": "user", "content": "x"}]) for s in sessions)
)
await asyncio.gather(*(s.get_items() for s in sessions))
parked = _count_open_fds()
assert parked is not None
# 60 parked sessions, yet descriptors are back at the baseline.
assert parked - baseline <= 5, f"parked fds grew by {parked - baseline}"
finally:
for s in sessions:
s.close()
@pytest.mark.asyncio
async def test_in_flight_descriptors_track_concurrency_not_session_count(
tmp_path: Path,
) -> None:
baseline = _count_open_fds()
if baseline is None:
pytest.skip("no /proc/self/fd or /dev/fd on this platform")
sessions = [open_agent_session(f"a{i}", tmp_path / f"s{i}.db") for i in range(200)]
peak = baseline
try:
async def sample() -> None:
nonlocal peak
for _ in range(500):
current = _count_open_fds()
if current is not None:
peak = max(peak, current)
await asyncio.sleep(0)
async def load() -> None:
for _ in range(4):
await asyncio.gather(
*(s.add_items([{"role": "user", "content": "x"}]) for s in sessions)
)
await asyncio.gather(load(), sample())
# 200 sessions, but peak is bounded by the thread pool, well under 200.
assert peak - baseline < 100, f"in-flight fds peaked at +{peak - baseline}"
finally:
for s in sessions:
s.close()
@pytest.mark.asyncio
async def test_history_survives_the_per_operation_connection(tmp_path: Path) -> None:
session = open_agent_session("agent-1", tmp_path / "agents.db")
try:
for i in range(30):
await session.add_items([{"role": "user", "content": f"m{i}"}])
items = [cast("dict[str, Any]", i) for i in await session.get_items()]
assert [i["content"] for i in items] == [f"m{i}" for i in range(30)]
finally:
session.close()
@pytest.mark.asyncio
async def test_concurrent_sessions_sharing_one_file_stay_consistent(tmp_path: Path) -> None:
db = tmp_path / "shared.db"
sessions = [open_agent_session(f"a{i}", db) for i in range(10)]
try:
await asyncio.gather(
*(s.add_items([{"role": "user", "content": s.session_id}]) for s in sessions)
)
# Each session sees only its own row despite sharing the file.
for s in sessions:
items = [cast("dict[str, Any]", i) for i in await s.get_items()]
assert [i["content"] for i in items] == [s.session_id]
finally:
for s in sessions:
s.close()
@pytest.mark.asyncio
async def test_a_closed_session_refuses_operations(tmp_path: Path) -> None:
session = open_agent_session("agent-1", tmp_path / "agents.db")
await session.add_items([{"role": "user", "content": "x"}])
session.close()
with pytest.raises(RuntimeError, match="closed"):
await session.add_items([{"role": "user", "content": "y"}])
+105
View File
@@ -0,0 +1,105 @@
from __future__ import annotations
import json
from typing import Any
import pytest
from agents.tool_context import ToolContext
from strix.tools.todo import tools
from strix.tools.todo.tools import _coerce_priority, create_todo
@pytest.fixture(autouse=True)
def _isolate_store() -> Any:
tools._todos_storage.clear()
yield
tools._todos_storage.clear()
async def _create(todos: list[Any], agent_id: str = "root") -> dict[str, Any]:
ctx = ToolContext(
context={"agent_id": agent_id},
tool_name="create_todo",
tool_call_id="call-1",
tool_arguments="{}",
)
raw = await create_todo.on_invoke_tool(ctx, json.dumps({"todos": json.dumps(todos)}))
return json.loads(raw) # type: ignore[no-any-return]
def test_unknown_priority_falls_back_to_normal() -> None:
assert _coerce_priority("medium") == "normal"
assert _coerce_priority("urgent") == "normal"
assert _coerce_priority("high") == "high"
@pytest.mark.asyncio
async def test_one_bad_priority_no_longer_discards_the_batch() -> None:
result = await _create(
[
{"title": "Recon", "priority": "medium"},
{"title": "Probe /admin", "priority": "sky-high"},
{"title": "Report"},
]
)
assert result["success"] is True
assert result["created_count"] == 3
by_title = {c["title"]: c["priority"] for c in result["created"]}
assert by_title["Recon"] == "normal"
assert by_title["Probe /admin"] == "normal"
assert by_title["Report"] == "normal"
@pytest.mark.asyncio
async def test_duplicate_titles_within_a_batch_are_skipped() -> None:
result = await _create(
[
{"title": "Subdomain enumeration"},
{"title": "Content discovery"},
{"title": "Subdomain enumeration"},
{"title": "content discovery"},
]
)
assert result["created_count"] == 2
assert {c["title"] for c in result["created"]} == {
"Subdomain enumeration",
"Content discovery",
}
assert len(result["skipped"]) == 2
assert all(s["reason"] == "duplicate title" for s in result["skipped"])
@pytest.mark.asyncio
async def test_a_title_already_on_the_list_is_not_created_again() -> None:
await _create([{"title": "Crawl with katana"}])
result = await _create([{"title": "crawl with katana"}, {"title": "JS analysis"}])
assert [c["title"] for c in result["created"]] == ["JS analysis"]
assert [s["title"] for s in result["skipped"]] == ["crawl with katana"]
assert result["total_count"] == 2
def test_coerce_never_raises() -> None:
assert _coerce_priority("nonsense") == "normal"
assert _coerce_priority(None) == "normal"
assert _coerce_priority("high") == "high"
for value in (2, ["high"], {"p": 1}, True):
assert _coerce_priority(value) == "normal" # type: ignore[arg-type]
@pytest.mark.asyncio
async def test_non_string_priority_does_not_fail_the_batch() -> None:
result = await _create(
[
{"title": "Recon", "priority": 2},
{"title": "Probe", "priority": ["high"]},
{"title": "Report"},
]
)
assert result["success"] is True
assert result["created_count"] == 3
assert {c["priority"] for c in result["created"]} == {"normal"}
+30 -11
View File
@@ -234,12 +234,11 @@ async def test_confirming_the_mount_starts_the_scan_without_a_target() -> None:
@pytest.mark.asyncio
async def test_declining_the_mount_returns_to_the_start_screen() -> None:
started = False
async def test_declining_the_mount_runs_without_one() -> None:
started: list[bool] = []
async def start(_verify: bool = True) -> None:
nonlocal started
started = True
async def start(verify: bool = True) -> None:
started.append(verify)
os.environ["STRIX_LLM"] = "anthropic/claude-sonnet-4"
os.environ["ANTHROPIC_API_KEY"] = "test-key"
@@ -250,14 +249,34 @@ async def test_declining_the_mount_returns_to_the_start_screen() -> None:
result = await controller.handle("setup.confirm_mount", {"approved": False})
assert result == {"approved": False}
# Nothing was prepared, so the session goes back to the start screen and can
# be launched again.
assert started is False
# Declining skips the directory; it does not abandon the scan.
assert started == [False]
assert controller.workspace_mount is None
assert controller.pending_workspace_mount is None
assert controller.setup_mode is True
assert controller.scan_started is False
assert controller.scan_state == "setup"
assert controller.setup_mode is False
assert controller.scan_started is True
assert controller.scan_state == "running"
@pytest.mark.asyncio
async def test_approving_the_mount_runs_with_it() -> None:
started: list[bool] = []
async def start(verify: bool = True) -> None:
started.append(verify)
os.environ["STRIX_LLM"] = "anthropic/claude-sonnet-4"
os.environ["ANTHROPIC_API_KEY"] = "test-key"
loader._cached = None
controller = TuiController(args(), on_start=start)
await controller.handle("setup.start", {"verify": False, "mount_working_dir": True})
result = await controller.handle("setup.confirm_mount", {"approved": True})
assert result == {"approved": True}
assert started == [False]
assert controller.workspace_mount == str(Path.cwd())
assert controller.scan_state == "running"
@pytest.mark.asyncio
+59 -3
View File
@@ -7,19 +7,26 @@ shows what the user actually typed; resuming has to match that.
from __future__ import annotations
import ast
import json
import sqlite3
from pathlib import Path
from typing import TYPE_CHECKING, Any
import pytest
from strix.core import execution
from strix.core.paths import runtime_state_dir
from strix.interface.tui.backend.live_view import TuiLiveView as GoTuiLiveView
from strix.interface.tui.live_view import TuiLiveView, _is_internal_agent_turn
from strix.interface.tui.live_view import (
_INTERNAL_TURN_PREFIXES,
TuiLiveView,
_is_internal_agent_turn,
)
if TYPE_CHECKING:
from pathlib import Path
from types import ModuleType
def _write_run(run_dir: Path, items: list[dict[str, Any]], agent_id: str = "root") -> None:
@@ -176,11 +183,60 @@ def test_internal_turn_classifier_matches_every_injected_form() -> None:
"[CRITICAL] Turn budget: 480/500 used (96%).",
"== Inherited context from parent (background only) ==",
"Your previous message ended a turn without a tool call.",
"Your previous response ended the autonomous Strix run without a lifecycle tool call.",
"Your previous response ended the autonomous run without a lifecycle tool call.",
):
assert _is_internal_agent_turn(content), content
def _injected_strings(module: ModuleType) -> list[str]:
"""Every string a module can inject, and nothing it merely mentions.
Parsing rather than searching the text keeps comments out of it, so a stale
copy of a message left in a comment cannot pass for the message itself. It
also joins adjacent literals for free, which the line wrapping needs, and
docstrings are dropped because they describe the code rather than run in it.
"""
tree = ast.parse(Path(module.__file__ or "").read_text(encoding="utf-8"))
docstrings = set()
for node in ast.walk(tree):
if not isinstance(node, ast.Module | ast.ClassDef | ast.FunctionDef | ast.AsyncFunctionDef):
continue
first = node.body[0] if node.body else None
if isinstance(first, ast.Expr) and isinstance(first.value, ast.Constant):
docstrings.add(id(first.value))
literals: list[str] = []
for node in ast.walk(tree):
if isinstance(node, ast.Constant):
if isinstance(node.value, str) and id(node) not in docstrings:
literals.append(node.value)
elif isinstance(node, ast.JoinedStr):
literals.append(
"".join(
part.value
for part in node.values
if isinstance(part, ast.Constant) and isinstance(part.value, str)
)
)
return literals
def test_internal_turn_prefixes_still_match_what_is_injected() -> None:
"""The classifier copies sentences out of another module, so they can drift.
Both nudges are written inline in strix.core.execution, so there is nothing to
import and compare against. Read them back out of what that module can inject.
"""
injected = _injected_strings(execution)
nudges = [prefix for prefix in _INTERNAL_TURN_PREFIXES if prefix.startswith("Your previous")]
assert nudges, "the no-tool-call nudges are no longer in the classifier"
for nudge in nudges:
assert any(nudge in literal for literal in injected), (
f"the classifier expects {nudge!r}, which strix.core.execution no longer "
f"injects. A resumed scan would show that nudge as the user's own message."
)
def test_internal_turn_classifier_keeps_bracketed_user_text() -> None:
"""A leading bracket is not enough: typed text often starts with one."""
for content in (
Generated
+1 -1
View File
@@ -2378,7 +2378,7 @@ wheels = [
[[package]]
name = "strix-agent"
version = "1.5.0"
version = "1.5.2"
source = { editable = "." }
dependencies = [
{ name = "caido-sdk-client" },