Compare commits

..
Author SHA1 Message Date
bearsyankees f5a467fca5 Add reverse-engineering security skills 2026-08-19 12:06:11 -04:00
31 changed files with 962 additions and 599 deletions
-8
View File
@@ -15,14 +15,6 @@ npx skills add usestrix/strix
- `fix-security-vulnerabilities-with-strix` — remediate findings and re-run Strix to verify - `fix-security-vulnerabilities-with-strix` — remediate findings and re-run Strix to verify
- `ci-security-scanning-with-strix` — add PR scanning to CI/CD (self-hosted CLI or managed app) - `ci-security-scanning-with-strix` — add PR scanning to CI/CD (self-hosted CLI or managed app)
Target-specific workflows built on the same engine:
- `application-security-testing` — whole-product AppSec review: pick the right test per asset, then rank the results
- `web-app-penetration-testing` — black-box pentest of a live web app or staging site
- `api-security-testing` — REST/GraphQL APIs and the OWASP API Security Top 10 (BOLA/IDOR, authz)
- `owasp-top-10-testing` — systematic OWASP Top 10 assessment with honest per-category coverage
- `find-security-vulnerabilities-in-code` — white-box review of a repo or working tree
**Two ways to run, same engine — pick per situation:** **Two ways to run, same engine — pick per situation:**
- **Open-source CLI (self-hosted):** free, fully local, BYO LLM key, needs Docker. Best for local dev loops, air-gapped/offline, and full control. - **Open-source CLI (self-hosted):** free, fully local, BYO LLM key, needs Docker. Best for local dev loops, air-gapped/offline, and full control.
+1 -6
View File
@@ -116,7 +116,7 @@ Strix is agent-ready. Give Claude Code, Cursor, Codex, or any [SKILL.md-compatib
npx skills add usestrix/strix npx skills add usestrix/strix
``` ```
This installs nine skills: **penetration-testing-with-strix** (run headless scans and read results), **managed-pentesting-with-strix** (drive the managed [app.strix.ai](https://app.strix.ai) platform via REST — no local Docker or LLM key), **fix-security-vulnerabilities-with-strix** (remediate + re-scan to verify), **ci-security-scanning-with-strix** (PR scanning in CI), plus target-specific workflows: **application-security-testing**, **web-app-penetration-testing**, **api-security-testing**, **owasp-top-10-testing**, and **find-security-vulnerabilities-in-code**. Agents can run Strix two ways with the same engine — the open-source CLI locally, or the managed cloud when there's no local infra — and read [`AGENTS.md`](AGENTS.md) for a quick reference, [docs.strix.ai/llms.txt](https://docs.strix.ai/llms.txt) for the CLI docs, and [docs.app.strix.ai](https://docs.app.strix.ai) for the API. This installs four skills: **penetration-testing-with-strix** (run headless scans and read results), **managed-pentesting-with-strix** (drive the managed [app.strix.ai](https://app.strix.ai) platform via REST — no local Docker or LLM key), **fix-security-vulnerabilities-with-strix** (remediate + re-scan to verify), and **ci-security-scanning-with-strix** (PR scanning in CI). Agents can run Strix two ways with the same engine — the open-source CLI locally, or the managed cloud when there's no local infra — and read [`AGENTS.md`](AGENTS.md) for a quick reference, [docs.strix.ai/llms.txt](https://docs.strix.ai/llms.txt) for the CLI docs, and [docs.app.strix.ai](https://docs.app.strix.ai) for the API.
--- ---
@@ -167,15 +167,10 @@ strix view
# ...or open a specific run by name # ...or open a specific run by name
strix view my-run-name strix view my-run-name
# Expose the viewer on all IPv4 interfaces at a fixed port
strix view --host 0.0.0.0 --port 8080 --no-open
``` ```
`strix view` starts a lightweight local server (bound to `127.0.0.1` on a random port) and opens your browser to a private, tokened link. Nothing leaves your machine: the dashboard reads the run's files straight off disk, with no cloud account or upload required. The UI ships prebuilt with Strix, so there is no extra install and no JS build step. `strix view` starts a lightweight local server (bound to `127.0.0.1` on a random port) and opens your browser to a private, tokened link. Nothing leaves your machine: the dashboard reads the run's files straight off disk, with no cloud account or upload required. The UI ships prebuilt with Strix, so there is no extra install and no JS build step.
Use `--host 0.0.0.0` to make the viewer reachable from other machines. Replace `0.0.0.0` in the printed URL with the server's reachable IP or hostname. The token in that URL grants access to the selected run's scan data, history, and steering, so only share it with trusted users and restrict the port with your firewall. Requests without the token-derived session cannot read run data.
### What's in the dashboard ### What's in the dashboard
- **Overview**: run status, target, and a severity breakdown of everything found so far. - **Overview**: run status, target, and a severity breakdown of everything found so far.
-5
View File
@@ -19,11 +19,6 @@ npx skills add usestrix/strix
| `managed-pentesting-with-strix` | Drive the managed [app.strix.ai](https://app.strix.ai) platform over REST — no local Docker or LLM key needed | | `managed-pentesting-with-strix` | Drive the managed [app.strix.ai](https://app.strix.ai) platform over REST — no local Docker or LLM key needed |
| `fix-security-vulnerabilities-with-strix` | Triage findings, fix root causes, and re-run Strix to verify each fix | | `fix-security-vulnerabilities-with-strix` | Triage findings, fix root causes, and re-run Strix to verify each fix |
| `ci-security-scanning-with-strix` | Add PR security scanning to GitHub Actions or any CI (self-hosted CLI or managed app) | | `ci-security-scanning-with-strix` | Add PR security scanning to GitHub Actions or any CI (self-hosted CLI or managed app) |
| `application-security-testing` | Assess a whole product: choose the right test for each asset, then rank the findings into one remediation plan |
| `web-app-penetration-testing` | Black-box pentest of a live web app or staging site — scope, credentials, and multi-account access-control testing |
| `api-security-testing` | Test a REST/GraphQL API against the OWASP API Security Top 10 — schema-driven enumeration, BOLA/IDOR, authz |
| `owasp-top-10-testing` | Systematic OWASP Top 10 assessment with honest per-category coverage |
| `find-security-vulnerabilities-in-code` | White-box security review of a repo or working tree, with exploits to confirm findings |
Install a single skill with `npx skills add usestrix/strix --skill penetration-testing-with-strix`, or use one without installing: Install a single skill with `npx skills add usestrix/strix --skill penetration-testing-with-strix`, or use one without installing:
-61
View File
@@ -1,61 +0,0 @@
---
name: api-security-testing
description: Security-test a REST, GraphQL, or gRPC API with Strix — autonomous agents that enumerate endpoints from an OpenAPI/GraphQL schema (or by crawling), then actually exploit the API-specific vulnerability classes in the OWASP API Security Top 10 (2023) — broken object-level authorization (BOLA/IDOR), broken object property level authorization (excessive data exposure and mass assignment), broken function-level authorization, unrestricted resource consumption, SSRF, injection, and auth/token flaws. Every finding comes with a working proof-of-concept request. Use when the user asks to pentest, security-test, audit, or find vulnerabilities in an API, endpoint, or backend service.
license: Apache-2.0
metadata:
author: usestrix
homepage: https://docs.strix.ai
---
# Security-test an API
APIs fail differently from web UIs: there is no rendered surface to crawl, the interesting bugs are authorization-shaped rather than injection-shaped, and the same endpoint behaves differently per token. This workflow targets those specifics with Strix's autonomous agents, using the current [OWASP API Security Top 10 (2023)](https://owasp.org/API-Security/editions/2023/en/0x11-t10/) as the coverage checklist. For the web-app equivalent, the current edition is the OWASP Top 10:2025 — see **owasp-top-10-testing**.
Install, LLM setup, full CLI flags, and the managed-cloud path are in the **penetration-testing-with-strix** skill. Read it if `strix --version` fails or the target is not an API.
## 1. Gather what the agents need
APIs are near-impossible to test blind, so collect first:
| Input | Why it matters |
|---|---|
| **Schema** — OpenAPI/Swagger file, Postman collection, GraphQL endpoint (introspection), or a gRPC `.proto` | Turns guesswork into full endpoint enumeration. Biggest single win in coverage. An OpenAPI/Swagger or Postman spec (`.json`/`.yaml`/`.yml`) is a target Strix takes directly; a `.proto` is not, so pass it with `--workspace-file`. |
| **Two sets of credentials/tokens**, ideally in different tenants | BOLA/IDOR — API1:2023, still the #1 API risk — can only be *proven* by accessing tenant A's objects with tenant B's token. |
| **A low-privilege and a high-privilege token** | Required to prove broken function-level authorization (API5:2023 — a `user` calling admin-only routes). |
| **Example object IDs** | Lets agents test ID tampering immediately instead of hunting for valid identifiers. |
| **Out-of-scope routes** | Payments, mass notification, destructive admin endpoints. |
| **Rate limits / WAF** in front of the API | Avoids agents burning budget on throttled requests; mention them so testing adapts. |
Ask the user for anything missing — do not fabricate tokens or scan an API they do not own.
## 2. Run the scan
Pass the spec as a **target**, not as prose in the instruction — Strix parses OpenAPI/Swagger (`.json`/`.yaml`) and Postman collection exports directly, so the agents start from the real endpoint list:
```bash
strix -n -t ./openapi.yaml -t https://api.staging.example.com --max-budget 20 \
--instruction "Tenant A token: <tokenA> (org 1111, user id 11, order id 501).
Tenant B token: <tokenB> (org 2222, user id 22).
Admin token: <tokenAdmin>.
Focus: BOLA across orgs (API1), function-level authz on /admin/* (API5), object property level authz on PATCH /users/{id} — both mass assignment and over-exposed fields in list responses (API3), unrestricted resource consumption (API4).
Out of scope: POST /billing/*, POST /notifications/broadcast."
```
- **Postman instead of OpenAPI:** a collection export works as a target (`-t ./collection.postman_collection.json`), or pull one live with `-t postman://<collection-uuid>` (optionally `"postman://<collection-uuid>?env=<environment-uuid>"`), which needs `POSTMAN_API_KEY` in the environment.
- **Many services at once:** put one target per line in a file and pass `--target-list ./targets.txt`, repeatable and combinable with `-t`.
- **Add the backend source for depth:** `-t ./services/api -t https://api.staging.example.com`. With code access the agents can reason about authorization checks and object ownership rather than inferring them from responses.
- **gRPC:** target the endpoint and pass the definition as a workspace file, `-t https://grpc.staging.example.com --workspace-file ./service.proto`. Only `.json`, `.yaml`, and `.yml` specs are recognized as targets, so `-t ./service.proto` fails with "Path exists but is not a directory".
- **GraphQL:** point at the GraphQL endpoint and say whether introspection is enabled; call out that you want batching/aliasing abuse, depth/complexity limits, and per-field authorization tested.
- **Internal/private APIs** unreachable from your machine: use the managed platform's network connector — see **managed-pentesting-with-strix**.
- Use `--instruction-file` when the credential/context block gets long, and keep tokens out of shell history and out of committed files.
- **Supporting files** the agents should read but not test, such as an endpoint wordlist or handwritten notes about the tenancy model: pass `--workspace-file ./notes.md`. The file lands read-only in `/workspace`. Add `:DEST` to choose the path, for example `--workspace-file ./wordlist.txt:lists/wordlist.txt`.
## 3. Verify findings
`strix_runs/<run>/penetration_test_report.md` first, then `vulnerabilities/*.md` — each contains the exact request that proved the issue. Replay it (for example, with `curl`) before reporting; for authorization findings, confirm the response really contains the other tenant's data rather than an empty 200.
`findings.sarif` uploads to GitHub code scanning; `vulnerabilities.json` is the structured index for ticketing.
## 4. Fix, re-test, and keep it tested
Remediate with **fix-security-vulnerabilities-with-strix** (fix the authorization check, not the single endpoint), then re-run against the same target to prove the exploit is dead. Wire it into pull-request CI with **ci-security-scanning-with-strix** so new endpoints get tested as they ship.
@@ -1,66 +0,0 @@
---
name: application-security-testing
description: Application security testing (AppSec) across a whole product with Strix — decide which asset needs which test (source code, running web app, API, CI pipeline), run it, and turn the results into a ranked remediation plan. Autonomous agents exploit and prove each issue instead of emitting static-analysis alerts, so the plan is ordered by what is actually reachable. Use when the user asks for an application security review or audit, an appsec assessment, vulnerability scanning across their stack, a security review before a launch or a customer security questionnaire, or does not yet know which kind of security test they need.
license: Apache-2.0
metadata:
author: usestrix
homepage: https://docs.strix.ai
---
# Application security testing
Entry point for "make my application secure" requests, where the target is not yet a single URL or repo. The job here is to pick the right test per asset, run it, and produce one ranked plan — not to run everything at maximum depth.
Install, LLM setup, all CLI flags, and the managed-cloud path live in the **penetration-testing-with-strix** skill. Read it first if `strix --version` fails.
Only test assets the user owns or is authorized to test. Confirm authorization before the first run, and prefer staging over production, because the agents send real exploit payloads and can change data.
## 1. Map the assets
Ask (or read from the repo) and write the answers down before scanning:
- **Source** — one repo, a monorepo, several services? Which languages/frameworks?
- **Running environments** — is there a staging deployment? A public production site? A local dev server only?
- **APIs** — REST, GraphQL, gRPC? Is there an OpenAPI/GraphQL schema?
- **Authentication** — can you get two test accounts in different tenants? Most high-impact bugs need them.
- **Constraints** — out-of-scope paths, whether production may be touched, budget and wall-clock limits.
If there is no staging environment and production is off limits, say so early. A code-only review is still valuable, but it cannot prove exploitability against a live app.
## 2. Pick the right test per asset
| Asset | Skill to use |
| --- | --- |
| Repository or working tree | **find-security-vulnerabilities-in-code** |
| Live web app or staging site | **web-app-penetration-testing** |
| REST/GraphQL/gRPC API | **api-security-testing** |
| Assessment mapped to OWASP categories | **owasp-top-10-testing** |
| Every pull request, continuously | **ci-security-scanning-with-strix** |
| No Docker, no LLM key, or a report an auditor will accept | **managed-pentesting-with-strix** |
Those skills carry the flags, credential handling, and result-reading details. Do not duplicate their instructions here.
Sequence for a first assessment:
1. Review the code. It is the cheapest run and it maps the authorization model.
2. Pentest staging with credentials, and pass the repo as a second target so the agents keep source context.
3. Add CI scanning, so later regressions are caught without another manual pass.
Run one asset at a time and read each report before starting the next. Findings from the code review make the live run sharper.
## 3. Consolidate into one plan
Findings arrive per run in `strix_runs/<run>/`. Merge them into a single list and rank by **proven impact**, not by scanner severity:
1. Validated exploits reachable without authentication.
2. Validated cross-tenant or privilege-escalation issues.
3. Validated issues needing an authenticated account.
4. Unproven observations (configuration, dependency, and hardening notes) — flag as such, and never present them as confirmed vulnerabilities.
Deduplicate: the same root cause often surfaces in both the code review and the live pentest.
## 4. Be honest about coverage
State plainly what was *not* tested — assets with no staging environment, categories a black-box run cannot reach (logging and alerting, supply-chain integrity, insecure design), and any run that hit its budget or turn cap before finishing. Check `run.json` status and cost against `--max-budget` for each run. An empty result set from a truncated scan is not a clean bill of health.
Then remediate with **fix-security-vulnerabilities-with-strix**, which re-runs Strix against each fix to prove the exploit no longer works.
@@ -12,7 +12,7 @@ metadata:
You can gate PRs two ways — pick based on the environment, or combine them: You can gate PRs two ways — pick based on the environment, or combine them:
- **Managed platform (recommended for most teams)** — connect the GitHub/GitLab/Bitbucket app once and Strix reviews every PR with **no workflow file, no runner, no Docker, and no LLM key**. Results post as PR comments and land in the team dashboard. Best when you want zero CI maintenance, central tracking, or your runners lack Docker. See "Managed platform" below and the **managed-pentesting-with-strix** skill. - **Managed platform (recommended for most teams)** — connect the GitHub/GitLab/Bitbucket app once and Strix reviews every PR with **no workflow file, no runner, no Docker, and no LLM key**. Results post as PR comments and land in the team dashboard. Best when you want zero CI maintenance, central tracking, or your runners lack Docker. See "Managed platform" below and the **managed-pentesting-with-strix** skill.
- **Self-hosted OSS CLI in your runner** — run a diff-scoped scan as a pipeline step. Fully in your infra, free (BYO LLM key), no external account. Requires Docker on the runner. Best for air-gapped/self-hosted CI or when you do not want scans leaving your environment. - **Self-hosted OSS CLI in your runner** — run a diff-scoped scan as a pipeline step. Fully in your infra, free (BYO LLM key), no external account. Requires Docker on the runner. Best for air-gapped/self-hosted CI or when you don't want scans leaving your environment.
Both fail the build on validated findings and both emit SARIF 2.1.0, so you can start with one and add the other later. Both fail the build on validated findings and both emit SARIF 2.1.0, so you can start with one and add the other later.
@@ -63,13 +63,13 @@ jobs:
fi fi
``` ```
Then tell the user to add two repository secrets: `STRIX_LLM` (model id, for example `openai/gpt-5.4`) and `LLM_API_KEY` (the provider key). Do not create these values yourself. Then tell the user to add two repository secrets: `STRIX_LLM` (model id, e.g. `openai/gpt-5.4`) and `LLM_API_KEY` (the provider key). Do not create these values yourself.
Notes: Notes:
- In CI/headless runs Strix automatically scopes to the PR's changed files (`--scope-mode auto`). If diff resolution fails, keep `fetch-depth: 0` or set `--diff-base` to the PR's actual base branch — use `origin/${{ github.base_ref }}` in GitHub Actions rather than a hard-coded `origin/main`, since repos use different default branches. - In CI/headless runs Strix automatically scopes to the PR's changed files (`--scope-mode auto`). If diff resolution fails, keep `fetch-depth: 0` or set `--diff-base` to the PR's actual base branch — use `origin/${{ github.base_ref }}` in GitHub Actions rather than a hard-coded `origin/main`, since repos use different default branches.
- Exit codes: `0` pass, `2` vulnerabilities found (fails the job), `1` setup error. - Exit codes: `0` pass, `2` vulnerabilities found (fails the job), `1` setup error.
- The runner needs Docker (default GitHub-hosted Ubuntu runners have it). - The runner needs Docker (default GitHub-hosted Ubuntu runners have it).
- **Size the budget so the scan completes — do not let it fail open.** A `0` exit means "no validated vulnerabilities in what was analyzed"; if `--max-budget` is hit before the diff is fully covered, the scan wraps up early and can still exit `0`. The "Fail unless the scan completed" step above narrows the gap: `strix_runs/<run>/run.json` is `"stopped"` when the scan was cut off at the hard budget limit without a final report. It is not a complete guard — the agents get graduated wrap-up warnings before that limit, and a run that wraps up on a warning still calls `finish_scan` and records `"completed"` with partial coverage. So keep that step in any pipeline that gates merges **and** give the scan real headroom (compare `run.json`'s `llm_usage.cost` against `--max-budget`; if it ran right up to the cap, raise it). For a `quick` diff-scoped PR scan `--max-budget 10` is usually ample, raise it for large diffs. - **Size the budget so the scan completes — don't let it fail open.** A `0` exit means "no validated vulnerabilities in what was analyzed"; if `--max-budget` is hit before the diff is fully covered, the scan wraps up early and can still exit `0`. The "Fail unless the scan completed" step above narrows the gap: `strix_runs/<run>/run.json` is `"stopped"` when the scan was cut off at the hard budget limit without a final report. It is not a complete guard — the agents get graduated wrap-up warnings before that limit, and a run that wraps up on a warning still calls `finish_scan` and records `"completed"` with partial coverage. So keep that step in any pipeline that gates merges **and** give the scan real headroom (compare `run.json`'s `llm_usage.cost` against `--max-budget`; if it ran right up to the cap, raise it). For a `quick` diff-scoped PR scan `--max-budget 10` is usually ample, raise it for large diffs.
### Optional: upload findings to GitHub code scanning ### Optional: upload findings to GitHub code scanning
@@ -90,7 +90,7 @@ Any pipeline works the same way — install, set the two env vars, run headless:
```bash ```bash
curl -sSL https://strix.ai/install | bash curl -sSL https://strix.ai/install | bash
# Resolve the PR's base branch robustly (use your CI's base-branch variable if it # Resolve the PR's base branch robustly (use your CI's base-branch variable if it
# has one, for example GitHub Actions: origin/${{ github.base_ref }}). Avoid piping the # has one, e.g. GitHub Actions: origin/${{ github.base_ref }}). Avoid piping the
# git lookup into another command — a failed lookup would otherwise be masked. # git lookup into another command — a failed lookup would otherwise be masked.
BASE_BRANCH="${CI_MERGE_REQUEST_TARGET_BRANCH_NAME:-}" # GitLab MR target BASE_BRANCH="${CI_MERGE_REQUEST_TARGET_BRANCH_NAME:-}" # GitLab MR target
if [ -z "$BASE_BRANCH" ]; then if [ -z "$BASE_BRANCH" ]; then
@@ -98,7 +98,7 @@ if [ -z "$BASE_BRANCH" ]; then
BASE_BRANCH="${BASE_BRANCH#origin/}" BASE_BRANCH="${BASE_BRANCH#origin/}"
fi fi
DIFF_BASE="origin/${BASE_BRANCH:-main}" DIFF_BASE="origin/${BASE_BRANCH:-main}"
# Fail loudly rather than silently narrowing scope (for example, to HEAD~1, which on a # Fail loudly rather than silently narrowing scope (e.g. to HEAD~1, which on a
# multi-commit branch would scan only the last commit and let earlier ones pass). # multi-commit branch would scan only the last commit and let earlier ones pass).
if ! git rev-parse --verify --quiet "$DIFF_BASE" >/dev/null; then if ! git rev-parse --verify --quiet "$DIFF_BASE" >/dev/null; then
echo "Cannot resolve diff base '$DIFF_BASE'. Fetch the base branch (git fetch origin <base>) or set --diff-base explicitly." >&2 echo "Cannot resolve diff base '$DIFF_BASE'. Fetch the base branch (git fetch origin <base>) or set --diff-base explicitly." >&2
@@ -1,62 +0,0 @@
---
name: find-security-vulnerabilities-in-code
description: Find security vulnerabilities in a codebase or repository with Strix — a white-box AI security review that reads your source, reasons about the actual data flow and authorization model, then exploits what it finds in a live sandbox so every reported issue has a working proof-of-concept instead of a noisy static-analysis alert. Covers injection, XSS, SSRF, broken access control and IDOR, insecure deserialization, secrets in code, unsafe dependencies, and business-logic flaws. Use when the user asks to security-scan, security-review, or audit their code, repo, or pull request for vulnerabilities.
license: Apache-2.0
metadata:
author: usestrix
homepage: https://docs.strix.ai
---
# Find security vulnerabilities in code
White-box security review with Strix: the agents read the source to build a model of routes, sinks, and authorization checks, then attempt real exploitation. Findings come with a proof-of-concept, so the output is a short list of proven issues rather than the hundreds of "potential" hits a pattern-matching scanner produces.
Install, LLM setup, all flags, and the managed-cloud path are in the **penetration-testing-with-strix** skill.
## Run it
```bash
# Local working tree
strix -n -t ./ --scan-mode standard --max-budget 15
# A GitHub repo directly
strix -n -t https://github.com/org/app --max-budget 15
# Monorepo: point at the service that matters, not the whole tree
strix -n -t ./services/checkout --max-budget 20
# Only what a branch changed (whole-repo review is wasteful on a large repo)
strix -n -t ./ --scope-mode diff --diff-base origin/main --max-budget 10
```
A local path is mounted into the sandbox **writable**, so the agents can modify it. Run against a clean checkout.
Two things sharply improve results:
1. **Add a running instance of the app.** `-t ./ -t http://host.docker.internal:3000` lets the agents confirm exploitability against live behavior instead of reasoning about it statically — this is the difference between "this looks unsafe" and a validated finding. If nothing is running, static-only findings should be described as unconfirmed.
2. **Scope the review.** Point at the risky subtree and say what matters:
```bash
strix -n -t ./services/api --max-budget 15 \
--instruction "Focus on the authorization layer in src/auth and every route under src/routes/admin. Multi-tenant app: tenant id comes from the JWT. Flag any query that filters by object id without also filtering by tenant."
```
Tenancy model, trust boundaries, and which inputs are attacker-controlled are things the agents cannot infer reliably — tell them.
## Reviewing a pull request instead of the whole repo
For diff-scoped review of a branch or PR (and blocking merges on findings), use **ci-security-scanning-with-strix** — it covers diff scoping, PR comments, and SARIF upload to GitHub code scanning. The managed platform can also review PRs directly via API (**managed-pentesting-with-strix**).
## Read the results
In `strix_runs/<run>/`: `penetration_test_report.md` (start here), `vulnerabilities/*.md` (one per finding, with PoC and remediation), `vulnerabilities.json` / `.csv`, `findings.sarif` (upload to code scanning), `run.json`.
Before reporting to the user, open each finding and check the PoC actually demonstrates impact. Report file and line alongside the exploit so the fix is obvious.
Exit `0` means nothing exploitable was proven in what was analyzed — not that the codebase is clean. Check `run.json` status and cost against `--max-budget`, and note which paths went unreviewed if the run was capped.
## Complementary tooling
This is exploit-validated review, not an exhaustive inventory. Keep a dependency scanner (SCA) and secret scanning in place for complete coverage of known-CVE dependencies and committed credentials; use this for the logic, authorization, and injection bugs those tools structurally cannot find.
## Fix and verify
Hand results to **fix-security-vulnerabilities-with-strix**: patch the root cause (the shared authorization helper, not the one route), then re-run Strix to prove the exploit no longer works.
@@ -27,7 +27,7 @@ Order work by severity: critical → high → medium → low. Every Strix findin
For each finding: For each finding:
1. Reproduce it with the PoC from the finding file when feasible. 1. Reproduce it with the PoC from the finding file when feasible.
2. Fix the root cause, not the specific payload (parameterize every query instead of blocking one string, and enforce authorization in the handler instead of hiding the endpoint). 2. Fix the root cause, not the specific payload (e.g. parameterize all queries, don't blocklist one string; enforce authorization in the handler, don't hide the endpoint).
3. Prefer the framework's built-in defense (ORM parameterization, template auto-escaping, CSRF middleware, centralized authz) over ad-hoc sanitization. 3. Prefer the framework's built-in defense (ORM parameterization, template auto-escaping, CSRF middleware, centralized authz) over ad-hoc sanitization.
4. Keep the diff minimal and apply the repo's existing patterns. Finding files often include `fix_before`/`fix_after` snippets — use them as a starting point, not verbatim. 4. Keep the diff minimal and apply the repo's existing patterns. Finding files often include `fix_before`/`fix_after` snippets — use them as a starting point, not verbatim.
@@ -70,7 +70,7 @@ new_id=$(curl -sS "$BASE/scans/$scan_id/rerun" "${auth[@]}" -X POST | jq -r .sca
Or, if the cloud scan came from a repo/PR, trigger a fresh PR review on the fix branch (`POST /pr-reviews/start`). The platform also retests a single finding directly: `POST /api/v1/vulnerabilities/{vulnerabilityId}/retest`. Or, if the cloud scan came from a repo/PR, trigger a fresh PR review on the fix branch (`POST /pr-reviews/start`). The platform also retests a single finding directly: `POST /api/v1/vulnerabilities/{vulnerabilityId}/retest`.
- Also re-run the PoC manually when it is a simple request/script — fastest signal. - Also re-run the PoC manually when it is a simple request/script — fastest signal.
- Run the project's own test suite to make sure the fix does not break behavior. - Run the project's own test suite to make sure the fix doesn't break behavior.
## 4. Report ## 4. Report
@@ -80,7 +80,7 @@ Useful `CreateScanRequest` fields:
| `domain_ids` / `repository_ids` / `internal_targets` | targets (at least one) | | `domain_ids` / `repository_ids` / `internal_targets` | targets (at least one) |
| `domain_paths` / `repository_branches` | narrow to specific paths / branches | | `domain_paths` / `repository_branches` | narrow to specific paths / branches |
| `credentials` | authenticated scanning, incl. `mfa_method` (`totp`/`email_otp`/…) + `totp_secret` | | `credentials` | authenticated scanning, incl. `mfa_method` (`totp`/`email_otp`/…) + `totp_secret` |
| `headers` | extra HTTP headers (API keys, for example) for the target | | `headers` | extra HTTP headers (e.g. API keys) for the target |
| `focus` / `concerns` / `context` | steer the agents | | `focus` / `concerns` / `context` | steer the agents |
| `upload_ids` | attach uploaded source/docs archives for white-box context | | `upload_ids` | attach uploaded source/docs archives for white-box context |
| `notify_on_completion` / `notification_emails` | email when done | | `notify_on_completion` / `notification_emails` | email when done |
@@ -89,7 +89,7 @@ Response is `{ scan_id, title, status }` with `status` = `pending`.
## 3. Poll to completion ## 3. Poll to completion
`GET /scans/{scanId}` (`scans:read`). Status flow: `pending → running → completed` (or `failed` / `cancelled`). Poll on an interval — scans take minutes to hours. Do not block. `GET /scans/{scanId}` (`scans:read`). Status flow: `pending → running → completed` (or `failed` / `cancelled`). Poll on an interval — scans take minutes to hours; don't block.
```bash ```bash
while :; do while :; do
@@ -143,10 +143,10 @@ List/inspect via `GET /pr-reviews` and `GET /pr-reviews/{id}`. Repo-level PR-rev
## 7. Continuous testing (schedules & webhooks) ## 7. Continuous testing (schedules & webhooks)
- **Schedules** (`schedules:write`, Pro plan): create recurring scans and trigger them on demand — the managed equivalent of a cron-driven CLI loop. - **Schedules** (`schedules:write`, Pro plan): create recurring scans and trigger them on demand — the managed equivalent of a cron-driven CLI loop.
- **Webhooks** (`webhooks:write`): subscribe to pentest/vulnerability lifecycle events such as `scan.completed` and `vulnerability.created` to push results into Slack, ticketing, or your own pipeline instead of polling. - **Webhooks** (`webhooks:write`): subscribe to pentest/vulnerability lifecycle events (e.g. `scan.completed`, `vulnerability.created`) to push results into Slack, ticketing, or your own pipeline instead of polling.
See the schedules and webhooks sections at [docs.app.strix.ai](https://docs.app.strix.ai) for payloads. See the schedules and webhooks sections at [docs.app.strix.ai](https://docs.app.strix.ai) for payloads.
## Safety ## Safety
Only scan assets the user's organization owns or is authorized to test. External domain scans require verification (DNS/file/meta-tag) enforced by the platform — do not try to bypass it. Only scan assets the user's organization owns or is authorized to test. External domain scans require verification (DNS/file/meta-tag) enforced by the platform — don't try to bypass it.
-64
View File
@@ -1,64 +0,0 @@
---
name: owasp-top-10-testing
description: Test an application against the OWASP Top 10 with Strix — autonomous AI agents that attempt real exploits for each category of the current OWASP Top 10:2025 (broken access control including SSRF, security misconfiguration, software supply chain failures, cryptographic failures, injection, insecure design, authentication failures, integrity failures, logging and alerting failures, mishandling of exceptional conditions) and report only what they could actually prove, mapped back to the category with a proof-of-concept. Also covers the OWASP API Security Top 10 (2023). Use when the user asks for an OWASP Top 10 assessment, OWASP compliance testing, or a security review mapped to OWASP categories.
license: Apache-2.0
metadata:
author: usestrix
homepage: https://docs.strix.ai
---
# Test against the OWASP Top 10
The OWASP Top 10 is a taxonomy of risk categories, not a test suite — "OWASP Top 10 testing" means exercising each category against the real application and reporting what's actually exploitable. Strix's agents do the exploitation; this skill covers running it category-by-category and reporting coverage honestly.
**Use the current edition: [OWASP Top 10:2025](https://owasp.org/Top10/)** (8th installment, superseding 2021). Ask the user before targeting an older edition — some compliance checklists still reference 2021, and a report labelled with the wrong edition is misleading. Key differences from 2021: **SSRF is folded into A01**, **A03 Software Supply Chain Failures** expands the old "Vulnerable and Outdated Components", and **A10 Mishandling of Exceptional Conditions** is new; A02 Security Misconfiguration moved 5→2.
Install, LLM setup, and the managed-cloud alternative: **penetration-testing-with-strix**.
## What is and is not testable by an agent
Be straight with the user about this — claiming a clean sweep of all ten is misleading.
| Category (2025) | Coverage |
|---|---|
| A01 Broken Access Control (incl. SSRF) | **Strong** — cross-user/tenant access, privilege escalation, IDOR, and SSRF (including blind, via out-of-band callbacks) are all exploit-validated. Needs two accounts plus a privileged one to prove the authorization half. |
| A02 Security Misconfiguration | **Strong** — debug endpoints, verbose errors, permissive CORS, missing hardening, default credentials, exposed admin surfaces. |
| A03 Software Supply Chain Failures | **Partial** — version fingerprinting, and vulnerable/outdated dependency review when source is supplied. Build-system and distribution-infrastructure compromise (the broader half of this category) is out of scope for a runtime scan — pair with SCA plus build-provenance controls. |
| A04 Cryptographic Failures | **Partial** — transport config, unencrypted data in transit, secrets and tokens leaked in responses. At-rest crypto and key management need source or infra review. |
| A05 Injection | **Strong** — SQL/NoSQL/command/template injection and XSS, exploit-validated. |
| A06 Insecure Design | **Partial** — business-logic abuse (price/quantity tampering, workflow skipping, race conditions) is found where reachable; design intent still needs human review and threat modelling. |
| A07 Authentication Failures | **Strong** — auth bypass, weak session/token handling, password-reset and MFA flaws. |
| A08 Software or Data Integrity Failures | **Partial** — insecure deserialization and unsigned-update paths where reachable; CI/CD trust boundaries are not runtime-testable. |
| A09 Security Logging & Alerting Failures | **Not testable from outside** — requires reviewing the logging and alerting pipeline. State this rather than reporting it as passed. |
| A10 Mishandling of Exceptional Conditions | **Partial** — agents actively probe error handling and fail-open behavior (malformed input, forced errors, race and timeout conditions) and report what leaks or bypasses a control; exhaustive coverage of internal error paths needs source review. |
For APIs, run the same exercise against the **OWASP API Security Top 10 (2023)** — API1 BOLA, API3 Broken Object Property Level Authorization (2019's excessive data exposure + mass assignment merged), API5 broken function-level authorization — using the **api-security-testing** skill.
## Run it
Maximum category coverage comes from giving the agents both the source and a running instance, plus credentials at two privilege levels:
```bash
strix -n \
-t https://github.com/org/app \
-t https://staging.example.com \
--scan-mode deep --max-budget 30 \
--instruction "OWASP Top 10:2025 assessment. Cover every category systematically and map each finding to its 2025 category id.
Accounts: userA@example.com/<pw> (org 1), userB@example.com/<pw> (org 2), admin@example.com/<pw>.
Prioritise A01 (cross-org access, privilege escalation, SSRF), A02, A05, A07, A10.
Out of scope: /billing/*, outbound email."
```
- `--scan-mode deep` matters here: systematically walking ten categories is not a quick scan.
- Without a second account, A01 results are structurally incomplete — say so in the report rather than leaving it implied.
- Need an auditor-facing PDF? Run it through the managed platform and pull the technical report (**managed-pentesting-with-strix**).
## Report honestly
From `strix_runs/<run>/`, group `vulnerabilities/*.md` by category and state, per category: what was attempted, what was proven, and what could not be assessed (A09 always; A03/A04/A06/A08/A10 partially). Label the report with the edition used. Verify each PoC yourself before it goes in front of the user.
A `0` exit code means nothing exploitable was proven **in what was analyzed** — check `run.json` status and cost against `--max-budget`; a budget-capped run is not a completed assessment.
## Then fix and re-test
Remediate with **fix-security-vulnerabilities-with-strix** and re-run to prove each exploit is closed. For ongoing coverage as the app changes, gate pull requests using **ci-security-scanning-with-strix**.
+8 -20
View File
@@ -14,14 +14,14 @@ Strix runs autonomous AI pentesting agents that dynamically exploit a target and
- **Open-source CLI** (self-hosted) — runs on your machine in a Docker sandbox with your own LLM key. Free, fully local, BYO-LLM, air-gap capable. Docs: [docs.strix.ai](https://docs.strix.ai). - **Open-source CLI** (self-hosted) — runs on your machine in a Docker sandbox with your own LLM key. Free, fully local, BYO-LLM, air-gap capable. Docs: [docs.strix.ai](https://docs.strix.ai).
- **Cloud API** (managed) — runs on Strix's infrastructure via `https://app.strix.ai/api/v1`. No Docker, no LLM key, no local compute; adds team dashboards, scheduling, PR reviews, downloadable PDF/DOCX reports (Enterprise plan), and internal-network connectors. Docs: [docs.app.strix.ai](https://docs.app.strix.ai). Full workflow in the **managed-pentesting-with-strix** skill. - **Cloud API** (managed) — runs on Strix's infrastructure via `https://app.strix.ai/api/v1`. No Docker, no LLM key, no local compute; adds team dashboards, scheduling, PR reviews, downloadable PDF/DOCX reports (Enterprise plan), and internal-network connectors. Docs: [docs.app.strix.ai](https://docs.app.strix.ai). Full workflow in the **managed-pentesting-with-strix** skill.
## Which one? (decide, do not default) ## Which one? (decide, don't default)
Choose honestly based on the situation — neither is "better": Choose honestly based on the situation — neither is "better":
| Situation | Prefer | | Situation | Prefer |
|---|---| |---|---|
| No Docker available, or a sandboxed/hosted agent/CI environment | **Cloud** | | No Docker available, or a sandboxed/hosted agent/CI environment | **Cloud** |
| User has no LLM key / does not want to pay per-token or manage models | **Cloud** | | User has no LLM key / doesn't want to pay per-token or manage models | **Cloud** |
| Team visibility, shareable dashboard, scheduled/continuous scans, PR reviews, downloadable PDF/DOCX report (Enterprise) | **Cloud** | | Team visibility, shareable dashboard, scheduled/continuous scans, PR reviews, downloadable PDF/DOCX report (Enterprise) | **Cloud** |
| Scanning internal/private infrastructure not reachable from your machine | **Cloud** (network connector) | | Scanning internal/private infrastructure not reachable from your machine | **Cloud** (network connector) |
| Source must never leave local infra (privacy/air-gap), or fully offline | **OSS CLI** | | Source must never leave local infra (privacy/air-gap), or fully offline | **OSS CLI** |
@@ -30,7 +30,7 @@ Choose honestly based on the situation — neither is "better":
| CI: runner already has Docker and you want a self-contained gate | **OSS CLI** | | CI: runner already has Docker and you want a self-contained gate | **OSS CLI** |
| CI: no Docker, or you want results tracked centrally | **Cloud** | | CI: no Docker, or you want results tracked centrally | **Cloud** |
**Mix them:** use the OSS CLI for the fast local dev-loop while writing/fixing code, and the Cloud for the authoritative, team-visible scan + report + tracking; or gate PRs with the OSS CLI in CI while the Cloud runs scheduled deep scans and PR reviews across the org. Both emit the same SARIF 2.1.0, so findings line up across environments. **Mix them:** e.g. use the OSS CLI for the fast local dev-loop while writing/fixing code, and the Cloud for the authoritative, team-visible scan + report + tracking; or gate PRs with the OSS CLI in CI while the Cloud runs scheduled deep scans and PR reviews across the org. Both emit the same SARIF 2.1.0, so findings line up across environments.
If unsure and the user has (or will create) an app.strix.ai account, prefer **Cloud** — it avoids all local-infra friction. If they want zero signup / full local control, use the **OSS CLI**. If unsure and the user has (or will create) an app.strix.ai account, prefer **Cloud** — it avoids all local-infra friction. If they want zero signup / full local control, use the **OSS CLI**.
@@ -70,33 +70,21 @@ strix -n -t https://github.com/org/app -t https://staging.example.com
strix -n -t https://app.example.com \ strix -n -t https://app.example.com \
--instruction "Use credentials user@example.com:pass123. Focus on IDOR and auth bypass." --instruction "Use credentials user@example.com:pass123. Focus on IDOR and auth bypass."
# API spec as a first-class target (OpenAPI/Swagger or a Postman collection export) # Large monorepo: bind-mount instead of copying
strix -n -t ./openapi.yaml -t https://api.staging.example.com strix -n --mount ./huge-monorepo
# Many targets from a file, one per line
strix -n --target-list ./targets.txt --max-budget 30
# Give the agents a file to work with (wordlist, spec, notes) without making it a target
strix -n -t https://staging.example.com --workspace-file ./wordlist.txt --max-budget 20
``` ```
A local path passed with `-t` is mounted into the sandbox **writable** — the agents can read and modify it, so point at a clean checkout, not uncommitted work you care about.
Key flags: Key flags:
| Flag | Meaning | | Flag | Meaning |
|---|---| |---|---|
| `-t, --target` | URL, repo URL, local path, domain, IP, OpenAPI/Postman spec, or `postman://<uuid>`. Repeatable. | | `-t, --target` | URL, repo URL, local path, domain, or IP. Repeatable. |
| `--target-list PATH` | File of targets, one per line (`#` comments allowed). Repeatable, combines with `-t`. |
| `-n, --non-interactive` | Headless, exits on completion. Required for agents. | | `-n, --non-interactive` | Headless, exits on completion. Required for agents. |
| `-m, --scan-mode` | `quick` (minutes) / `standard` (~30 min) / `deep` (hours, default). | | `-m, --scan-mode` | `quick` (minutes) / `standard` (~30 min) / `deep` (hours, default). |
| `--instruction` / `--instruction-file` | Credentials, focus areas, scope rules. | | `--instruction` / `--instruction-file` | Credentials, focus areas, scope rules. |
| `--workspace-file PATH[:DEST]` | Place a file from this machine into `/workspace` read-only before the scan, for a wordlist, a spec, or notes. Repeatable. |
| `--max-budget USD` | Hard LLM spend cap; scan wraps up cleanly at the limit. | | `--max-budget USD` | Hard LLM spend cap; scan wraps up cleanly at the limit. |
| `--max-turns N` | Per-agent turn cap (default 500). | | `--max-turns N` | Per-agent turn cap (default 500). |
| `--resume RUN_NAME` | Resume a prior run from `strix_runs/`, with its agent history and targets. Cannot be combined with `-t`. | | `--resume RUN_NAME` | Resume a prior run from `strix_runs/`. |
| `--scope-mode` | For code targets: `auto` (diff-scope in CI/headless), `diff` (force changed files only), `full` (whole tree). |
| `--diff-base REF` | Branch or commit that `diff` scope compares against. Defaults to the repo's default branch. |
Scans take minutes (`quick`) to hours (`deep`). Run them in the background and poll for completion rather than blocking. Scans take minutes (`quick`) to hours (`deep`). Run them in the background and poll for completion rather than blocking.
@@ -142,7 +130,7 @@ curl -sS "$BASE/scans/$scan_id" -H "Authorization: Bearer $STRIX_API_TOKEN" | jq
curl -sS "$BASE/scans/$scan_id/sarif" -H "Authorization: Bearer $STRIX_API_TOKEN" -o findings.sarif curl -sS "$BASE/scans/$scan_id/sarif" -H "Authorization: Bearer $STRIX_API_TOKEN" -o findings.sarif
``` ```
Ask the user to create the token (and register the target as a domain/repository asset) if they have not. If Docker/local prerequisites are not already satisfied, use this path instead of trying to install infra. Ask the user to create the token (and register the target as a domain/repository asset) if they haven't. If Docker/local prerequisites aren't already satisfied, use this path instead of trying to install infra.
--- ---
@@ -1,54 +0,0 @@
---
name: web-app-penetration-testing
description: Pentest a web app or website end to end — black-box testing of a live URL, staging environment, or local dev server that finds and exploits real vulnerabilities (auth bypass, broken access control, IDOR, injection, XSS, SSRF, business logic) and proves each one with a working proof-of-concept instead of a signature match. Runs with Strix, either the self-hosted open-source CLI or the managed app.strix.ai cloud. Use when the user asks to pentest, hack, security-test, or audit their web app, website, web application, or staging site.
license: Apache-2.0
metadata:
author: usestrix
homepage: https://docs.strix.ai
---
# Pentest a web application
Black-box (and optionally source-assisted) penetration testing of a running web app with Strix's autonomous agents. Every reported finding is validated with a working exploit, so there are no signature-based false positives to triage.
Install, LLM setup, all CLI flags, and the managed-cloud alternative are covered in the **penetration-testing-with-strix** skill — read it if the target is not a running web app, or if `strix --version` fails. This skill is the web-app-specific workflow.
## 1. Confirm authorization and scope
Before running anything, establish:
- **The target is the user's** (or they are explicitly authorized to test it). Never pentest a third-party site on a hunch.
- **Which environment.** Prefer staging over production; agents send real exploit payloads and will create/modify data.
- **Out-of-scope paths** — payment flows, mass-email endpoints, admin destructive actions, third-party SSO providers.
- **Credentials.** Most real vulnerabilities live behind login. Without a test account, the agents only ever see the marketing surface.
Ask for anything missing rather than guessing.
## 2. Run the scan
```bash
strix -n -t https://staging.example.com --max-budget 20 \
--instruction "Test account: qa@example.com / <password>. In scope: /app/*, /api/*. Do not touch /billing or send email. Focus on access control between the two seeded orgs."
```
Notes that matter for web apps specifically:
- **Give it credentials via `--instruction`** (or `--instruction-file` for anything long), including how to log in if the flow is unusual (magic link, SSO, MFA-exempt test user).
- **Two accounts beat one.** Multi-tenant IDOR and broken-access-control bugs — consistently the highest-impact class in web apps — can only be proven when the agent can attempt cross-account access.
- **Add the repo for white-box depth** when you have the source: `-t https://github.com/org/app -t https://staging.example.com` (or a local path). Source access materially improves coverage of business-logic and authorization flaws.
- **Localhost works.** Point at `http://host.docker.internal:3000` (Docker Desktop) so the sandbox can reach a dev server on the host.
- `--scan-mode quick` for a fast dev-loop pass, `standard` (~30 min) for a normal review, `deep` for pre-release assurance. Always set `--max-budget`.
For a hosted run with no Docker/LLM key, or when the user wants a shareable dashboard and an auditor-ready PDF, use the cloud path in **managed-pentesting-with-strix** instead — same engine, same findings.
## 3. Review results
Read `strix_runs/<run>/penetration_test_report.md` first, then per-finding files in `vulnerabilities/`. Each contains the PoC — re-run it yourself to confirm before reporting to the user.
Exit codes: `0` no validated vulns in what was analyzed, `2` vulnerabilities found, `1` fatal error. A `0` is not proof of full coverage — if the budget or turn cap was hit the scan wraps up early, so check `run.json` status and cost against `--max-budget` before calling the app clean.
## 4. Fix and verify
Hand findings to the **fix-security-vulnerabilities-with-strix** skill: patch the root cause, then re-run Strix against the same target to prove the exploit no longer works. Re-testing is the only reliable confirmation a fix landed.
To keep the app tested on every change rather than once, wire Strix into CI with **ci-security-scanning-with-strix**.
+10 -47
View File
@@ -2,7 +2,6 @@
from __future__ import annotations from __future__ import annotations
import dataclasses
import inspect import inspect
import json import json
import logging import logging
@@ -223,17 +222,6 @@ def _with_coerced_arguments(tool: FunctionTool) -> FunctionTool:
return tool return tool
def _with_strictness(tool: FunctionTool, strict_schemas: bool) -> FunctionTool:
"""Drop strict JSON-schema mode when the route can't take it (see
``supports_strict_tool_schemas``); the tool stays functionally identical.
Returns a copy so the shared tool singletons keep their declared mode.
"""
if strict_schemas or not tool.strict_json_schema:
return tool
return dataclasses.replace(tool, strict_json_schema=False)
def _function_tool_with_error_result(tool: FunctionTool) -> FunctionTool: def _function_tool_with_error_result(tool: FunctionTool) -> FunctionTool:
invoke_tool = tool.on_invoke_tool invoke_tool = tool.on_invoke_tool
@@ -297,38 +285,24 @@ def _bound_custom_tool(tool: CustomTool) -> CustomTool:
return tool return tool
def _configure_filesystem_tools( def _configure_filesystem_tools(toolset: Any, *, chat_completions: bool) -> None:
toolset: Any, *, chat_completions: bool, strict_schemas: bool = True
) -> None:
for name, tool in vars(toolset).items(): for name, tool in vars(toolset).items():
if chat_completions: if chat_completions:
if isinstance(tool, CustomTool): if isinstance(tool, CustomTool):
setattr(toolset, name, _custom_tool_as_function_tool(tool)) setattr(toolset, name, _custom_tool_as_function_tool(tool))
elif isinstance(tool, FunctionTool): elif isinstance(tool, FunctionTool):
setattr( setattr(
toolset, toolset, name, _function_tool_with_error_result(_with_coerced_arguments(tool))
name,
_function_tool_with_error_result(
_with_strictness(_with_coerced_arguments(tool), strict_schemas)
),
) )
elif isinstance(tool, CustomTool): elif isinstance(tool, CustomTool):
setattr(toolset, name, _bound_custom_tool(tool)) setattr(toolset, name, _bound_custom_tool(tool))
elif isinstance(tool, FunctionTool): elif isinstance(tool, FunctionTool):
setattr( setattr(toolset, name, _with_bounded_result(_with_coerced_arguments(tool)))
toolset,
name,
_with_bounded_result(
_with_strictness(_with_coerced_arguments(tool), strict_schemas)
),
)
def _make_filesystem_configurator(*, chat_completions: bool, strict_schemas: bool) -> Any: def _make_filesystem_configurator(*, chat_completions: bool) -> Any:
def configure(toolset: Any) -> None: def configure(toolset: Any) -> None:
_configure_filesystem_tools( _configure_filesystem_tools(toolset, chat_completions=chat_completions)
toolset, chat_completions=chat_completions, strict_schemas=strict_schemas
)
return configure return configure
@@ -432,13 +406,11 @@ def _wrap_write_stdin(tool: FunctionTool) -> FunctionTool:
return tool return tool
def _configure_shell_tools( def _configure_shell_tools(toolset: Any, *, chat_completions: bool) -> None:
toolset: Any, *, chat_completions: bool, strict_schemas: bool = True
) -> None:
for name, tool in vars(toolset).items(): for name, tool in vars(toolset).items():
if not isinstance(tool, FunctionTool): if not isinstance(tool, FunctionTool):
continue continue
wrapped = _with_strictness(_with_coerced_arguments(tool), strict_schemas) wrapped = _with_coerced_arguments(tool)
if tool.name == "exec_command": if tool.name == "exec_command":
wrapped = _wrap_exec_command(wrapped) wrapped = _wrap_exec_command(wrapped)
elif tool.name == "write_stdin": elif tool.name == "write_stdin":
@@ -448,11 +420,9 @@ def _configure_shell_tools(
setattr(toolset, name, wrapped) setattr(toolset, name, wrapped)
def _make_shell_configurator(*, chat_completions: bool, strict_schemas: bool) -> Any: def _make_shell_configurator(*, chat_completions: bool) -> Any:
def configure(toolset: Any) -> None: def configure(toolset: Any) -> None:
_configure_shell_tools( _configure_shell_tools(toolset, chat_completions=chat_completions)
toolset, chat_completions=chat_completions, strict_schemas=strict_schemas
)
return configure return configure
@@ -598,7 +568,6 @@ def build_strix_agent(
is_whitebox: bool = False, is_whitebox: bool = False,
interactive: bool = False, interactive: bool = False,
chat_completions_tools: bool = False, chat_completions_tools: bool = False,
strict_tool_schemas: bool = True,
system_prompt_context: dict[str, Any] | None = None, system_prompt_context: dict[str, Any] | None = None,
extra_tools: Sequence[Tool] | None = None, extra_tools: Sequence[Tool] | None = None,
instructions_override: str | None = None, instructions_override: str | None = None,
@@ -608,8 +577,6 @@ def build_strix_agent(
Args: Args:
chat_completions_tools: Wrap SDK custom tools as function tools chat_completions_tools: Wrap SDK custom tools as function tools
when the selected backend cannot accept Responses custom tools. when the selected backend cannot accept Responses custom tools.
strict_tool_schemas: Send function tools as strict-schema tools. Off
for routes that reject a toolset this size as strict.
extra_tools: Additional tools for this scan agent only, on top of any extra_tools: Additional tools for this scan agent only, on top of any
registered via ``register_agent_tools``. registered via ``register_agent_tools``.
instructions_override: Use this verbatim as the system prompt instead instructions_override: Use this verbatim as the system prompt instead
@@ -637,7 +604,7 @@ def build_strix_agent(
tools = [*_BASE_TOOLS, *agent_tools, agent_finish] tools = [*_BASE_TOOLS, *agent_tools, agent_finish]
_ensure_unique_tool_names(tools) _ensure_unique_tool_names(tools)
tools = [ tools = [
_with_bounded_result(_with_strictness(_with_coerced_arguments(tool), strict_tool_schemas)) _with_bounded_result(_with_coerced_arguments(tool))
if isinstance(tool, FunctionTool) if isinstance(tool, FunctionTool)
else tool else tool
for tool in tools for tool in tools
@@ -663,13 +630,11 @@ def build_strix_agent(
Filesystem( Filesystem(
configure_tools=_make_filesystem_configurator( configure_tools=_make_filesystem_configurator(
chat_completions=chat_completions_tools, chat_completions=chat_completions_tools,
strict_schemas=strict_tool_schemas,
), ),
), ),
Shell( Shell(
configure_tools=_make_shell_configurator( configure_tools=_make_shell_configurator(
chat_completions=chat_completions_tools, chat_completions=chat_completions_tools,
strict_schemas=strict_tool_schemas,
), ),
), ),
], ],
@@ -682,7 +647,6 @@ def make_child_factory(
is_whitebox: bool = False, is_whitebox: bool = False,
interactive: bool = False, interactive: bool = False,
chat_completions_tools: bool = False, chat_completions_tools: bool = False,
strict_tool_schemas: bool = True,
system_prompt_context: dict[str, Any] | None = None, system_prompt_context: dict[str, Any] | None = None,
) -> Any: ) -> Any:
"""Return the runner-owned builder used by ``spawn_child_agent``. """Return the runner-owned builder used by ``spawn_child_agent``.
@@ -701,7 +665,6 @@ def make_child_factory(
is_whitebox=is_whitebox, is_whitebox=is_whitebox,
interactive=interactive, interactive=interactive,
chat_completions_tools=chat_completions_tools, chat_completions_tools=chat_completions_tools,
strict_tool_schemas=strict_tool_schemas,
system_prompt_context=system_prompt_context, system_prompt_context=system_prompt_context,
) )
-15
View File
@@ -749,18 +749,6 @@ def uses_chat_completions_tool_schema(model_name: str, settings: Settings) -> bo
return not model_supports_reasoning(model_name) return not model_supports_reasoning(model_name)
def supports_strict_tool_schemas(model_name: str) -> bool:
"""Return whether the route accepts strict tool schemas for Strix's toolset.
Claude caps a request at 20 strict tools and 16 union-typed parameters
across all strict schemas. Strix ships ~30 tools and the strict dialect
turns every optional parameter into a nullable union, so both caps are
exceeded and the request is rejected outright.
"""
name = model_name.strip().lower()
return not any(marker in name for marker in _ANTHROPIC_MODEL_MARKERS)
def model_supports_reasoning(model_name: str) -> bool: def model_supports_reasoning(model_name: str) -> bool:
import litellm import litellm
@@ -857,9 +845,6 @@ def is_known_openai_bare_model(model_name: str) -> bool:
return bool(entry and entry.get("litellm_provider") == "openai") return bool(entry and entry.get("litellm_provider") == "openai")
_ANTHROPIC_MODEL_MARKERS = ("anthropic", "claude", "sonnet", "opus", "haiku")
def is_claude_model(model_name: str) -> bool: def is_claude_model(model_name: str) -> bool:
return "claude" in (model_name or "").strip().lower() return "claude" in (model_name or "").strip().lower()
-6
View File
@@ -22,7 +22,6 @@ from strix.config import load_settings
from strix.config.models import ( from strix.config.models import (
StrixProvider, StrixProvider,
configure_sdk_model_defaults, configure_sdk_model_defaults,
supports_strict_tool_schemas,
uses_chat_completions_tool_schema, uses_chat_completions_tool_schema,
) )
from strix.config.settings import DEFAULT_MAX_TURNS from strix.config.settings import DEFAULT_MAX_TURNS
@@ -176,9 +175,6 @@ async def run_strix_scan(
) )
logger.info("LLM model resolved: %s", resolved_model) logger.info("LLM model resolved: %s", resolved_model)
chat_completions_tools = uses_chat_completions_tool_schema(resolved_model, settings) chat_completions_tools = uses_chat_completions_tool_schema(resolved_model, settings)
strict_tool_schemas = supports_strict_tool_schemas(resolved_model)
if not strict_tool_schemas:
logger.info("Sending non-strict tool schemas: %s caps strict tools", resolved_model)
if coordinator is None: if coordinator is None:
coordinator = AgentCoordinator() coordinator = AgentCoordinator()
@@ -310,7 +306,6 @@ async def run_strix_scan(
is_whitebox=is_whitebox, is_whitebox=is_whitebox,
interactive=interactive, interactive=interactive,
chat_completions_tools=chat_completions_tools, chat_completions_tools=chat_completions_tools,
strict_tool_schemas=strict_tool_schemas,
system_prompt_context=root_context, system_prompt_context=root_context,
instructions_override=root_instructions, instructions_override=root_instructions,
) )
@@ -329,7 +324,6 @@ async def run_strix_scan(
is_whitebox=is_whitebox, is_whitebox=is_whitebox,
interactive=interactive, interactive=interactive,
chat_completions_tools=chat_completions_tools, chat_completions_tools=chat_completions_tools,
strict_tool_schemas=strict_tool_schemas,
system_prompt_context=scope_context, system_prompt_context=scope_context,
) )
+1 -1
View File
@@ -346,7 +346,7 @@ def _load_resume_state(args: argparse.Namespace, parser: argparse.ArgumentParser
) )
try: try:
state = read_run_record(run_dir) state = read_run_record(run_dir)
except (RuntimeError, TypeError) as exc: except RuntimeError as exc:
parser.error(f"--resume {args.resume}: run.json unreadable: {exc}") parser.error(f"--resume {args.resume}: run.json unreadable: {exc}")
args.targets_info = state.get("targets_info") or [] args.targets_info = state.get("targets_info") or []
+2 -4
View File
@@ -146,9 +146,7 @@ def bounded_state_projection(state: dict[str, Any]) -> dict[str, Any]:
} }
for message in state["messages"][-5:] for message in state["messages"][-5:]
] ]
state["usage"] = { state["usage"] = {}
key: state["usage"][key] for key in ("total_tokens", "cost") if key in state["usage"]
}
state["error"] = terminal_projection(state["error"], max_string=512) state["error"] = terminal_projection(state["error"], max_string=512)
state["model_warning"] = terminal_projection(state["model_warning"], max_string=256) state["model_warning"] = terminal_projection(state["model_warning"], max_string=256)
state["caido_url"] = terminal_projection(state["caido_url"], max_string=256) state["caido_url"] = terminal_projection(state["caido_url"], max_string=256)
@@ -175,7 +173,7 @@ def bounded_state_projection(state: dict[str, Any]) -> dict[str, Any]:
"model_warning": "", "model_warning": "",
"caido_url": None, "caido_url": None,
"messages": [], "messages": [],
"usage": state["usage"], "usage": {},
"subscription": state["subscription"], "subscription": state["subscription"],
"viewer_status": state["viewer_status"], "viewer_status": state["viewer_status"],
"viewer_url": None, "viewer_url": None,
@@ -100,7 +100,7 @@ func applyMarkdownStyles(text string) string {
case strings.HasPrefix(line, "- "), strings.HasPrefix(line, "* "): case strings.HasPrefix(line, "- "), strings.HasPrefix(line, "* "):
out.WriteString(Col(Green).Render("• ") + inlineFormat(line[2:])) out.WriteString(Col(Green).Render("• ") + inlineFormat(line[2:]))
case len(line) > 2 && line[0] >= '0' && line[0] <= '9' && (line[1:3] == ". " || line[1:3] == ") "): case len(line) > 2 && line[0] >= '0' && line[0] <= '9' && (line[1:3] == ". " || line[1:3] == ") "):
out.WriteString(Col(Green).Render(line[:2]+" ") + inlineFormat(line[3:])) out.WriteString(Col(Green).Render(string(line[0])+". ") + inlineFormat(line[2:]))
case line == "---" || line == "***" || line == "___": case line == "---" || line == "***" || line == "___":
out.WriteString(Col(Green).Render(strings.Repeat("─", 40))) out.WriteString(Col(Green).Render(strings.Repeat("─", 40)))
default: default:
@@ -72,19 +72,6 @@ func TestNonTablePipeLinesAreLeftAlone(t *testing.T) {
} }
} }
func TestMarkdownOrderedListsUseSingleSpaceAfterMarker(t *testing.T) {
out := renderAssistantMarkdown("1. hello\n2) world")
plain := ansi.Strip(out)
for _, want := range []string{"1. hello", "2) world"} {
if !strings.Contains(plain, want) {
t.Fatalf("ordered list item %q missing: %q", want, plain)
}
}
if strings.Contains(plain, "1. hello") || strings.Contains(plain, "2) world") {
t.Fatalf("double space after the list marker: %q", plain)
}
}
func TestInlineFormatKeepsNonEmphasisMarkers(t *testing.T) { func TestInlineFormatKeepsNonEmphasisMarkers(t *testing.T) {
literal := []string{ literal := []string{
"ls *.py *.go", "ls *.py *.go",
+1 -5
View File
@@ -45,11 +45,7 @@ def run_view(argv: list[str]) -> None:
default=0, default=0,
help="Port to serve on (default: an available ephemeral port).", help="Port to serve on (default: an available ephemeral port).",
) )
parser.add_argument( parser.add_argument("--host", default="127.0.0.1", help=argparse.SUPPRESS)
"--host",
default="127.0.0.1",
help="Host to bind to (default: 127.0.0.1; use 0.0.0.0 for all IPv4 interfaces).",
)
parser.add_argument( parser.add_argument(
"--no-open", "--no-open",
action="store_true", action="store_true",
+18 -20
View File
@@ -135,9 +135,8 @@ class _ViewerState:
# exchanged for a session cookie only when presented on the initial page # exchanged for a session cookie only when presented on the initial page
# load. It is the request-level authorization the review asked for: # load. It is the request-level authorization the review asked for:
# reachability of the port (e.g. when bound with ``--host``) is not # reachability of the port (e.g. when bound with ``--host``) is not
# enough to read run data, steer a live scan, trigger a report, or # enough to steer a live scan, trigger a report, or browse history --
# browse history -- the token is never handed to a caller who merely # the token is never handed to a caller who merely reaches ``/``.
# reaches ``/``.
self.session_token = secrets.token_urlsafe(32) self.session_token = secrets.token_urlsafe(32)
# Finalized in ``serve()`` once the port is known (the server binds # Finalized in ``serve()`` once the port is known (the server binds
# after this state is constructed); see SESSION_COOKIE_PREFIX. # after this state is constructed); see SESSION_COOKIE_PREFIX.
@@ -235,11 +234,11 @@ def _make_handler(state: _ViewerState) -> type[BaseHTTPRequestHandler]:
self.end_headers() self.end_headers()
def _handle_api(self, path: str, query: dict[str, list[str]]) -> None: def _handle_api(self, path: str, query: dict[str, list[str]]) -> None:
# The cross-run history list (/api/runs) unlocks its entries only for # The launched run is always viewable with no verification. The
# a caller that holds this process's session capability *and* is # cross-run history list (/api/runs) unlocks its entries only for a
# email verified, so merely reaching an exposed --host port never # caller that holds this process's session capability *and* is email
# leaks the run list (the payload still advertises the count as a # verified, so merely reaching an exposed --host port never leaks the
# teaser). # run list (the payload still advertises the count as a teaser).
if path == "/api/runs": if path == "/api/runs":
unlocked = self._has_session() and auth.is_verified() unlocked = self._has_session() and auth.is_verified()
payload = build_runs_payload(state.base_dir, verified=unlocked) payload = build_runs_payload(state.base_dir, verified=unlocked)
@@ -254,13 +253,6 @@ def _make_handler(state: _ViewerState) -> type[BaseHTTPRequestHandler]:
self._handle_auth_status() self._handle_auth_status()
return return
# All remaining GET endpoints expose run metadata or scan output.
# Require the capability even for the run used to launch the viewer;
# reachability of an exposed --host port must not grant data access.
if not self._has_session():
self._send_json(HTTPStatus.FORBIDDEN, {"error": "forbidden"})
return
run_values = query.get("run") run_values = query.get("run")
run_param = run_values[0] if run_values else None run_param = run_values[0] if run_values else None
run_dir = resolve_run_dir(state.base_dir, run_param, state.run_dir) run_dir = resolve_run_dir(state.base_dir, run_param, state.run_dir)
@@ -268,10 +260,16 @@ def _make_handler(state: _ViewerState) -> type[BaseHTTPRequestHandler]:
self._send_json(HTTPStatus.NOT_FOUND, {"error": "unknown run"}) self._send_json(HTTPStatus.NOT_FOUND, {"error": "unknown run"})
return return
# Any run other than the one used to launch the viewer is part of the # The launched run is always viewable. Any *other* run's data is part
# email-gated history. The session check above applies to both paths; # of the gated history: it needs this process's session capability
# verification adds a second gate for historical run data. # (so merely reaching an exposed --host port is not enough) *and*
if run_dir.resolve() != state.run_dir.resolve() and not auth.is_verified(): # email verification -- otherwise knowing a run name would leak its
# metadata, vulnerabilities, report, and transcript.
if run_dir.resolve() != state.run_dir.resolve():
if not self._has_session():
self._send_json(HTTPStatus.FORBIDDEN, {"error": "forbidden"})
return
if not auth.is_verified():
self._send_json(HTTPStatus.UNAUTHORIZED, {"error": "unverified"}) self._send_json(HTTPStatus.UNAUTHORIZED, {"error": "unverified"})
return return
@@ -387,7 +385,7 @@ def _make_handler(state: _ViewerState) -> type[BaseHTTPRequestHandler]:
except auth.RelayError as exc: except auth.RelayError as exc:
self._send_relay_error(exc) self._send_relay_error(exc)
return return
# The password is returned only to a session-authorized browser. # The password is returned only to the local (127.0.0.1) browser.
self._send_json( self._send_json(
HTTPStatus.OK, HTTPStatus.OK,
{"ok": True, "password": password, "filename": filename}, {"ok": True, "password": password, "filename": filename},
+6
View File
@@ -55,6 +55,12 @@ Notable LLM security skills:
- `llm_applications` (technologies): end-to-end OWASP 2026 LLM01-LLM10 coverage across models, RAG, vectors, agents, tools, outputs, supply chain, and resource controls - `llm_applications` (technologies): end-to-end OWASP 2026 LLM01-LLM10 coverage across models, RAG, vectors, agents, tools, outputs, supply chain, and resource controls
- `llm_prompt_injection` (vulnerabilities): deep direct, indirect, multimodal, memory, and tool-result prompt-injection testing - `llm_prompt_injection` (vulnerabilities): deep direct, indirect, multimodal, memory, and tool-result prompt-injection testing
Notable reverse-engineering skills:
- `advisory_to_poc` (custom): advisory-to-root-cause workflow for patch diffing, public PoCs, and detector design
- `appliance_firmware` (technologies): appliance artifact, runtime, and install-state analysis
- `protocol_reverse_engineering` (protocols): stateful/custom protocol reconstruction and controlled harnessing
- `memory_corruption` (vulnerabilities): native crash triage, primitive quality, and exploitability constraints
--- ---
## 🎨 Creating New Skills ## 🎨 Creating New Skills
+235
View File
@@ -0,0 +1,235 @@
---
name: advisory-to-poc
description: Vulnerability research workflow for turning advisories, patches, release artifacts, public PoCs, and incident clues into root-cause analysis, safe reproducers, reliable detectors, patch-bypass review, and adjacent-bug hypotheses
---
# Advisory to PoC
Use this skill for authorized product-security and n-day research where the starting point is an advisory, fixed release, patch, public PoC, or incident evidence rather than a known vulnerable endpoint.
The goal is a version-bounded root-cause explanation and reliable, reproducible validation. Do not equate a changed function, crash, scanner hit, or advisory claim with exploitability.
## Evidence Ledger
Keep facts, inferences, and experiments separate:
| Type | Examples |
|---|---|
| Published fact | affected versions, CWE, exposed feature, vendor mitigation |
| Artifact fact | changed function, new validation, removed route, configuration delta |
| Inference | likely attacker-controlled field, suspected auth path, probable sink |
| Experiment | vulnerable response, fixed response, crash, OAST callback, file canary |
Record source URL, artifact hash, product edition/branch, build number, platform, configuration, and date. Re-check assumptions whenever the experimental result conflicts with the advisory narrative.
## Research Workflow
### 1. Scope the Claim
- Extract affected and fixed versions, branches, platforms, roles, protocols, and feature/configuration prerequisites.
- Note whether the vendor describes impact, root cause, mitigation, or only a CWE category.
- Treat bundled CVEs and large release rollups as multiple candidate changes until proven otherwise.
- Identify whether the issue is pre-auth, low-privilege, post-auth, local, or requires a victim/session bridge.
### 2. Acquire Comparable Artifacts
Prefer the closest vulnerable/fixed pair for the same edition and platform:
- source commits, tags, tests, pull requests, and dependency lockfiles
- packages, containers, installers, JAR/WAR/DLL/assemblies, Python bytecode, firmware, or VM images
- web-server/reverse-proxy configuration, service definitions, scripts, and bundled third-party components
- documentation and shipped examples that reveal routes, protocols, defaults, or extension points
Hash originals and work on copies. Preserve installation lineage: default credentials, generated keys, legacy files, and retained configs may matter even if a fresh fixed install does not contain them.
### 3. Reduce Diff Noise
Start with inventories before line-by-line analysis:
- added/removed/renamed files and dependencies
- changed routes, authorization annotations, allowlists/denylists, parser calls, command construction, length checks, and deserialization types
- edge configuration changes that block or rewrite a route without changing application code
- tests added, removed, or updated; these often encode a near-ready reproducer
- sibling call sites of the changed helper or validator
For binaries, combine string/import/symbol diffing with a decompiler and a second diffing method when possible. Large compiler or bundled-library changes create false clusters; anchor on advisory-relevant constants, protocol handlers, response strings, and call graphs.
### 4. Map External Reachability
Work from both directions:
```text
external listener -> edge config -> router -> authentication -> parser -> sink
known changed sink -> callers -> route/protocol -> authentication -> external listener
```
Inventory auxiliary listeners, management agents, sidecars, localhost APIs, custom RPC services, CGI/script dispatch, and framework direct-component routes. Do not assume the main web UI's authentication protects every product service.
Record branch-specific and configuration-specific exposure. A powerful sink behind a disabled feature or unreachable route is not a pre-auth vulnerability.
### 5. Explain the Patch Mechanism
State what security invariant the patch tries to restore:
- bounds, termination, initialization, or length/type consistency
- authentication/authorization before dispatch
- canonicalization before comparison
- allowlisted deserialization or reflection targets
- safe command/process APIs instead of shell construction
- file path confinement and extension/handler restrictions
- route removal or edge blocking
- session-field filtering or trustworthy state reconstruction
Then ask what the patch did not change: alternate callers, sibling parsers, secondary routes, nested gadgets, transitive deserialization, old aliases, different protocol handlers, and edge/application disagreement.
### 6. Build a Reproducer Ladder
Escalate one capability at a time:
1. **Presence** - product/version/protocol fingerprint with low noise
2. **Reachability** - expected route/parser/handler responds
3. **Security differential** - unauthorized behavior differs from a denied control
4. **Primitive** - safe read, controlled callback, canary write, harmless constructor, or deterministic crash in an isolated lab
5. **Impact** - demonstrate the requested authorized impact and preserve its prerequisites
Prefer distinctive non-secret response structure, benign errors, OAST DNS/HTTP callbacks, inert file markers, or no-op commands. For deserialization, use a non-executing network gadget before command execution. For memory corruption, establish the bug and mitigation constraints in a lab; a connection close or crash is not proof of RCE.
### 7. Calibrate on Controls
Run the same reproducer against:
- vulnerable version
- fixed version
- unaffected neighboring version where available
- feature disabled / hardened configuration
- malformed but non-triggering negative input
- authentication present vs absent, if the claim crosses an auth boundary
Repeat enough times to distinguish deterministic behavior from crashes, timing noise, worker restarts, load balancers, and transient network failures.
### 8. Hunt Adjacent and Partial Fixes
After reproducing the primary issue:
- enumerate every call site of the patched function/validator
- cluster nearby handlers using the same parser, session format, command wrapper, or file primitive
- replay the old PoC and structural variants against the first fixed version
- inspect whether the patch blocks the route while leaving the sink reachable elsewhere
- test nested/transitive objects rather than only top-level denylisted types
- check whether one advisory/CVE bundles multiple distinct vulnerable paths
Do not call a variant a bypass until the fixed version demonstrably remains vulnerable.
## Tool Routing
Use the lightest maintained tool that answers the current question. Pin versions in research notes and preserve generated outputs so another analyst can reproduce the diff.
### Artifact and Package Diff: diffoscope
[diffoscope](https://diffoscope.org/) is the default first pass for packages, directories, archives, and binaries. Use it to build a changed-file/config/package manifest before opening a decompiler. For hostile artifacts, keep inputs read-only, disable network, and run the helper-heavy comparison in an isolated environment.
### Firmware and Appliance Artifacts
When the starting point is firmware, a virtual appliance, or a nested image format, load `appliance_firmware`. That skill owns extraction, package/rootfs/runtime correlation, Ghidra/BinDiff routing, overlay/install-state analysis, and device-lifecycle caveats.
### Java/JVM: Vineflower
Use maintained [Vineflower](https://github.com/Vineflower/vineflower) for JAR/class decompilation. Diff archive inventories before decompiled text; compiler, obfuscator, and synthetic-code changes produce noise. Confirm suspicious control flow with bytecode (`javap -c`) rather than treating reconstructed Java as source truth.
### .NET: ILSpy / ilspycmd
Use [ILSpy](https://github.com/icsharpcode/ILSpy) for managed assemblies. Work offline, inspect IL/metadata when the C# reconstruction is ambiguous, and use only GitHub Releases or NuGet.
### Native Code: Ghidra and BinDiff
Use official [Ghidra](https://github.com/NationalSecurityAgency/ghidra) for cross-architecture disassembly/decompilation and [BinDiff](https://github.com/google/bindiff) only after the file/package diff has narrowed the relevant binaries. Keep the toolchain pinned, offline where practical, and non-executing. Decompiler output and similarity scores are triage aids, not proof.
## Source and Binary Techniques
### Source-Available Products
- Search route declarations, filters/interceptors, auth decorators, and direct framework component dispatch.
- Trace attacker-controlled fields through type coercion, validation, shell/process APIs, filesystem operations, reflection, template/XSLT evaluation, and deserialization.
- Compare callers, not just the patched callee. The same helper may be safe in one route and exposed in another.
- Read tests and examples for expected protocol syntax and serialized message shapes.
### Managed Artifacts
- Decompile JAR/WAR and .NET assemblies; diff namespaces/classes/method bodies and embedded configuration.
- Trace public setters, opaque identifiers, type metadata, and framework serialization hooks.
- Inspect bundled libraries and version changes, but prove application reachability before assigning impact.
### Native Binaries and Firmware
- Inventory architecture, mitigations, imports, strings, services, and exposed ports before deep reversing.
- Diff functions around new bounds checks, initialization, string termination, length casts, command builders, and protocol parsers.
- Reconstruct the smallest valid protocol state machine before mutating the suspected field.
- Use debuggers, sanitizers, traces, and process monitors inside an isolated lab when available.
- Separate bug existence from exploitability under ASLR, NX, stack canaries, allocator behavior, architecture, and restart model.
### Public PoC or Incident First
- First decompose and neutralize a public or captured PoC; reproduce its stages in an isolated lab while preserving the headers, ordering, sessions, and negotiation relevant to each stage.
- Decompose the PoC into stages and identify the oracle for each stage.
- Work backward from the final sink to root cause and forward from the entry point to confirm reachability.
- If no patch pair exists, controlled honeypot/instrumentation can reveal in-the-wild request structure; never expose a live vulnerable system beyond an isolated, monitored environment.
Pair `protocol_reverse_engineering` when the external entry point is binary, TLS-wrapped, message-oriented, or stateful.
## Detector Design
A detector must distinguish the vulnerable behavior reliably from fixed and unaffected behavior:
- match a structural response or deterministic state change, not a secret value
- use a unique per-target canary and clean it up when the test writes data
- distinguish patched denial from generic 404/500, WAF blocking, authentication failure, and connection loss
- complete protocol/session prerequisites instead of relying on a single raw request
- rate-limit crash-prone or resource-intensive probes and keep them opt-in
- calibrate templates against vulnerable, fixed, and negative-control targets
When scaling, separate fingerprinting from exploitation. Presence can prioritize assets; it does not confirm the vulnerability.
## Exploitability Triage
Rate each condition explicitly:
- attacker position and credentials
- default vs optional feature/configuration
- internet-facing vs auxiliary/local listener
- data/byte/control precision
- restart, race, victim action, or environment requirements
- available mitigations and architecture
- reliable primitive vs crash-only or unstable behavior
- practical post-primitive chain in the product's default deployment
Down-rate unrealistic chains even when the underlying bug is real. Conversely, revisit “low” primitives such as SSRF, reflection, arbitrary write, cache control, or information disclosure in product context; native admin features may convert them into RCE.
## Validation Deliverable
Include:
1. exact affected/fixed artifacts and hashes
2. authoritative published claims and unresolved ambiguity
3. minimal relevant diff and restored invariant
4. external route/protocol and auth/config prerequisites
5. source-to-sink or packet-to-sink trace
6. safe reproducer plus positive and negative controls
7. vulnerable vs fixed results across repeat runs
8. exploitability constraints and why the demonstrated impact follows
9. adjacent paths reviewed and any partial-fix evidence
## Anti-Patterns
- Trusting the advisory CWE/title as the actual root cause
- Diffing only application code while ignoring edge/proxy/service configuration
- Treating any crash, close, 500, scanner alert, or changed function as exploitation
- Running a weaponized public PoC before isolating its stages and side effects
- Claiming pre-auth impact without tracing the complete auth and routing path
- Assuming one CVE maps to one code path or one patch fixes the whole vulnerability class
- Searching only for the published payload instead of the restored invariant
- Reporting a registry/download/callback signal without separating automated noise from authentic target execution
- Generalizing from one appliance/version/configuration without testing prerequisites
## Summary
Advisory-driven research is evidence-driven reverse engineering. Acquire comparable artifacts, reduce the diff to a security invariant, prove external reachability, climb a safe reproducer ladder, calibrate against fixed and negative controls, and then audit sibling paths and partial fixes. The reusable output is the method and invariant—not the vendor-specific exploit string.
@@ -0,0 +1,176 @@
---
name: protocol-reverse-engineering
description: Authorized analysis of undocumented, proprietary, binary, or stateful network protocols using passive captures, client/server artifacts, explicit state machines, bounded lab harnesses, and semantic vulnerable-versus-fixed validation
---
# Protocol Reverse Engineering
Use this skill when an exposed service cannot be tested correctly as isolated HTTP-like requests: custom RPC, binary framing, TLS-wrapped management protocols, message queues, VPN negotiation, in-band control records, or any protocol whose authentication and parsing depend on prior state.
The objective is a reviewable protocol model and controlled evidence that proves or disproves a security property. A socket connection, completed TLS handshake, `200`, or parser crash does not prove authentication, authorization, or code execution.
## Authorization and Safety Boundary
- Work from supplied artifacts, offline captures, or an isolated lab target unless active testing is explicitly authorized.
- Prefer offline parsing. Captures may contain credentials, session material, personal data, or private topology; minimize, encrypt, redact, and expire them.
- Never replay production credentials or captured authentication material.
- Put active harnesses in a network namespace or isolated VLAN with an explicit destination allowlist, low rate, bounded retries, and one mutation at a time.
- Do not broadcast, scan unrelated addresses, or start mutation/fuzz loops by default.
- Treat a malformed-packet crash as a denial-of-service test. Perform it only in a restartable lab and never infer RCE from it.
## Build the Protocol Model
Record each layer separately:
| Layer | Questions |
|---|---|
| Transport | TCP, UDP, HTTP tunnel, queue, Unix socket, reconnect behavior? |
| Security | TLS/mTLS, certificate role, message MAC/signature, encryption boundary? |
| Framing | magic, version, type, flags, length, checksum, terminator, nesting? |
| State | negotiation, challenge, authentication, session, command, teardown? |
| Identity | where is peer/user/device identity introduced and verified? |
| Authorization | which state or role permits each operation? |
| Data model | integers, strings, TLV, XML/JSON, compression, serialization? |
| Responses | acknowledgements, errors, correlation IDs, timing, connection close? |
Maintain a message-field ledger:
```text
offset/path | size/type | endian/encoding | producer | consumer | validation | state | confidence
```
Label every statement as observed, inferred, or experimentally confirmed. Unknown bytes remain unknown; do not name them after a single sample.
## Workflow
### 1. Collect Passive Evidence
Use, in order of preference:
- official protocol or integration documentation
- offline captures of a legitimate client/server exchange
- client binaries, SDKs, schemas, constants, error strings, and debug logs
- server handlers, dispatch tables, configuration, and certificate logic
- vulnerable/fixed captures or binaries from the same branch
Use two supplied or explicitly authorized successful sessions and controlled variations when available. Otherwise record the evidence gap; do not obtain or replay production credentials merely to complete the model. Compare message boundaries, counters, nonces, lengths, identity fields, and state-dependent responses. Keep the original capture immutable and hash it.
Use [TShark](https://www.wireshark.org/docs/man-pages/tshark.html) for reproducible offline extraction:
```bash
tshark -r session.pcapng -q -z conv,tcp
tshark -r session.pcapng -Y 'tcp.stream == 0' -T fields \
-e frame.number -e tcp.seq -e tcp.len -e tcp.payload
```
Prefer `-r` over live capture. Do not run Wireshark/TShark as root, capture unrelated production traffic, or assume dissector output is safe or correct; use a patched build in an isolated environment for hostile captures.
### 2. Reconstruct Framing Before Meaning
- Reassemble streams before assigning message boundaries; TCP packets are not application messages.
- Test length hypotheses against multiple messages and both directions.
- Identify byte order, signedness, alignment, padding, compression, and checksums.
- Separate outer transport/tunnel framing from the inner application message.
- For nested formats, model each parser boundary independently.
- Reject impossible lengths before allocation, recursion, decompression, or slicing.
When the layout stabilizes, encode it in a declarative grammar such as [Kaitai Struct](https://kaitai.io/). Add `valid` constraints and strict size/count limits; generated parsers can still allocate or recurse dangerously on hostile lengths. Keep compiler/runtime versions aligned and regression-test the grammar on positive, truncated, oversized, and unknown-type samples.
### 3. Recover the State Machine
Write transitions explicitly:
```text
DISCONNECTED -> TRANSPORT -> NEGOTIATED -> PEER_VERIFIED
-> USER_AUTHENTICATED -> AUTHORIZED -> OPERATION
```
For every transition, record:
- initiating message and required prior state
- server-side check and identity source
- success, denial, and malformed responses
- state stored across messages or reconnects
- timeout/replay/counter behavior
- whether an alternate message type reaches the same handler
Distinguish transport establishment, peer verification, user authentication, session creation, role authorization, and successful privileged action. Prove the specific boundary relevant to the security claim.
### 4. Trace Fields to Decisions and Sinks
From binaries or source, anchor on message IDs, error strings, constants, certificate handling, dispatcher tables, and changed functions. Trace attacker-controlled fields through:
- length arithmetic, allocation, copy, termination, and integer conversion
- parser state, tag nesting, recursion, and unknown-field behavior
- identity selection, trust flags, signature/certificate verification, and session lookup
- shell/process calls, filesystem paths, deserialization, reflection, or product-native admin operations
Decompiler output is a hypothesis. Confirm important conditions in assembly, bytecode, runtime logs, or controlled packet results.
### 5. Build a Bounded Active Harness
Only craft packets after valid framing and state are understood. [Scapy](https://scapy.readthedocs.io/en/stable/) is appropriate for packet layers and stateful automata:
```bash
python -m pip install 'scapy==<reviewed-version>'
```
Start with a local responder or replay parser, not the appliance. Preserve a known-good transcript, mutate one semantic field, recompute dependent lengths/checksums, and compare the response. The harness must enforce:
- exact destination/port allowlist
- one target and one mutation by default
- rate, packet count, response size, timeout, and retry ceilings
- no broadcast/multicast and no automatic crash retry
- artifact logging without credentials or secret payloads
- cleanup and target health check after each risky case
Raw sockets may require privilege; isolate socket creation and drop privileges afterward where possible.
### 6. Design Semantic Experiments
Prefer experiments that answer one question:
- Does an invalid identity or signature reach the authorized state?
- Does a declared length govern copying, parsing, or only framing?
- Do duplicate/unknown fields change the selected handler?
- Does patched behavior add validation, change state, or block an outer route?
- Does a response prove the operation, or merely that dispatch began?
Use vulnerable, fixed, and malformed-negative controls. Repeat enough to separate deterministic semantics from loss, retransmission, process restart, load balancing, and timeout noise.
## Safe Oracles
Prefer, from least to most invasive:
1. distinctive protocol/version field
2. deterministic denial-versus-accept response
3. synthetic-account no-op or non-secret lab read
4. unique constant callback through explicitly authorized, preferably self-hosted OAST
5. inert canary write with cleanup
6. process execution only under separate explicit authorization when no lower-harm oracle can establish the required impact
A connection close is normally an ambiguous result. If crash validation is unavoidable, combine lab-only process logs, restart evidence, and a non-triggering control; report bug existence separately from exploitability.
When the starting point is an advisory, fixed build, patch, or public PoC, pair this skill with `advisory_to_poc` for evidence classification, artifact comparison, and partial-fix review.
## Patch and Version Differentials
- Compare message/state behavior across the closest vulnerable and fixed builds of the same branch.
- Derive a fingerprint from the restored invariant, not only from banners.
- Check configuration, certificate role, feature enablement, architecture, and deployment mode.
- Treat protocol differences as version evidence unless they directly prove vulnerable behavior.
- When one handler is patched, enumerate sibling message types, alternate transports, and pre-auth dispatch paths using the same parser or decision.
## Validation Deliverable
Include:
1. target versions, platform, configuration, and artifact/capture hashes
2. layered protocol diagram and message-field ledger
3. explicit state machine and identity/authentication/authorization boundaries
4. source/binary trace for the relevant field and decision
5. bounded harness with rate/destination safeguards
6. vulnerable, fixed, and negative-control results
7. minimum safe oracle and any side effects/cleanup
8. unresolved fields, assumptions, and confidence levels
9. bug-existence versus exploitability assessment
@@ -0,0 +1,253 @@
---
name: appliance-firmware
description: Security analysis of appliances and firmware through artifact provenance, safe extraction, root filesystem and runtime mapping, listener and trust-boundary inventory, patch comparison, managed/native code triage, hardware constraints, and isolated device validation
---
# Appliance and Firmware Analysis
Use this skill for VPNs, firewalls, storage/backup systems, management appliances, embedded products, virtual appliances, and other packaged systems where security behavior is split across firmware, web-server configuration, native daemons, scripts, managed services, generated state, and hardware-specific runtime details.
Appliance research is architecture research. The public web UI is only one entry point; auxiliary listeners, localhost APIs, sidecars, support agents, update services, telemetry jobs, package installers, and product-native administration features often carry equal or greater authority.
## Build and Artifact Matrix
Record before comparing anything:
| Dimension | Examples |
|---|---|
| Product | model/SKU, physical/virtual/cloud image, edition/license |
| Software | marketing version, build/revision, branch, hotfix, package set |
| Platform | architecture, endian, kernel, libc, bootloader, filesystem |
| Install state | factory image, upgraded system, migrated config, retained files |
| Configuration | feature flags, listeners, authentication mode, HA/cluster role |
| Artifact source | vendor download, updater, installed disk, backup, marketplace |
| Update form | full image, delta package, component hotfix, rollback bundle |
| Authenticity | signature/encryption state, certificate/key ID, manifest/base-version requirement |
Hash original artifacts and preserve acquisition metadata. A neighboring version from a different SKU, edition, architecture, or installation lineage can produce a convincing but irrelevant diff.
## Safe Extraction
Treat firmware and every embedded archive/filesystem as hostile input. Extract as an unprivileged user into a fresh writable quota-limited output directory with no network, bounded recursion/processes, and read-only input.
### unblob
[unblob](https://github.com/onekey-sec/unblob) provides recursive extraction plus structured metadata for many firmware/container/filesystem formats. Prefer a reviewed container image digest:
```bash
appliance_out="$(mktemp -d)"
docker run --rm --network none \
--read-only --cap-drop ALL --security-opt no-new-privileges \
--user "$(id -u):$(id -g)" --pids-limit 256 --memory 4g --cpus 2 \
--tmpfs /tmp:rw,noexec,nosuid,size=512m \
-v /path/to/input:/data/input:ro \
-v "$appliance_out":/data/output \
ghcr.io/onekey-sec/unblob@sha256:<reviewed-digest> \
-e /data/output -d 6 -p 2 --report /data/output/unblob.json \
/data/input/firmware.bin
```
Create the output directory first and ensure it is writable by the chosen UID/GID; otherwise the host may create a root-owned mount point. Never extract over an existing analysis tree. Inspect symlinks, device nodes, archive paths, decompression ratios, and output size before interacting with the tree.
### diffoscope
Use [diffoscope](https://diffoscope.org/) for a recursive format-aware first comparison of vulnerable/fixed directories, packages, images, JARs, and executables:
```bash
diffoscope --html diffoscope.html vulnerable-root/ fixed-root/
```
Run it in an isolated reviewed container when processing hostile artifacts because it invokes many external format helpers. Use the first report to narrow files/config/packages rather than repeatedly expanding the entire image.
Use the unblob report and packaged filesystem metadata for ownership, mode, xattr, capability, and device-node claims; a host extraction run under your own UID can intentionally remap them. Do not mount an untrusted extracted filesystem or `chroot` into it on the analyst host.
## Filesystem and Boot Architecture
Inventory:
- partition table, bootloader, kernel, initramfs, SquashFS/UBIFS/ext filesystems
- init system, service definitions, inetd/socket activation, rc scripts, supervisors, and watchdogs
- read-only base image versus writable overlay, tmpfs, bind mounts, containers/chroots, and persistent data partitions
- factory defaults, first-boot generation, upgrade/migration scripts, rollback slots, and retained legacy files
- environment files, credentials, certificates, secrets, licenses, databases, sessions, caches, and backup/restore formats
- cron/timers, log rotation, telemetry, diagnostics, update checks, package deployment, support bundles, and cleanup tasks
- ownership, group membership, capabilities, setuid/setgid, ACLs, sudo/doas rules, device access, and IPC permissions
Static extracted files may not match runtime. Boot-time scripts can patch files, mount overlays, generate configs, copy certificates, activate routes, or replace binaries. Capture live filesystem/mount/process state when an apparently relevant change is absent from the disk image.
## Update and Installed-State Reconstruction
Before trusting a package or image diff, reconstruct how the device installs it:
- verify signature and manifest order, trust anchors, and whether integrity/authenticity checks cover the whole payload or only a wrapper
- distinguish full image, delta update, component hotfix, and required base version
- identify target partition, boot slot, rollback path, and anti-rollback/version checks
- review pre/post-install hooks, migrations, symlink changes, permission/capability changes, and retained/generated state
- map overlay, bind-mount, and generated-file precedence over the extracted rootfs
- test fresh install versus upgraded and partially rolled-back states
- reconcile package contents with hashes/build IDs from the actual running process and live filesystem
Record package-manager databases, shipped SBOM/manifests, bundled library copies, loader path, and `RPATH`/`RUNPATH` so you can distinguish a vulnerable library on disk from the library the running process actually maps.
## Listener and Service Map
Build a table for every network and local endpoint:
```text
address/port/socket | transport/TLS | process | config/init source
route/message type | authentication | authorization | privilege | feature/default
```
Include:
- HTTP(S) UI/API, CGI/FastCGI, WebSocket, SOAP, SAML/OIDC, upload/download
- SSH/SFTP, VPN/IKE, message queues, databases, backup/storage protocols
- proprietary TLS/RPC, cluster/HA, device-manager, agent, and telemetry ports
- loopback/Unix sockets, localhost APIs, sidecars, containers, and debug/support agents
- outbound update/download endpoints and trusted remote control planes
For outbound updater, telemetry, licensing, or control-plane names, record authoritative DNS/ownership, TLS identity and pinning, proxy/fallback behavior, request data, failure behavior, manifest integrity, payload integrity, rollback/version policy, and whether the external domain, bucket, package, or provider resource can expire or be reassigned.
Map edge configuration to code: reverse-proxy rules, rewrites, location blocks, authentication modules, trusted client-IP headers, TLS client certificates, and backend socket selection. A handler can be patched while a new edge rule merely hides it—or vice versa.
## Trust and Authorization Boundaries
Trace:
```text
external listener -> proxy/config -> router/dispatcher -> authentication
-> parser -> privileged operation -> OS/service identity
```
Test conceptual boundaries such as:
- public versus management interface
- external versus localhost/sidecar trust
- managed device versus manager/controller trust
- cluster peer, certificate, flag, or registration state
- web user versus OS/service/database authentication
- direct route versus internal redirect/component dispatch
- fresh install versus upgraded/retained installation state
- optional feature disabled versus installed-but-reachable handler
Successful TCP/TLS/WebSocket negotiation proves transport reachability, not authenticated identity or authorization. Determine the actual privileged result and which server-side flag/session/role enabled it.
## Code and Configuration Triage
### Scripts and Configuration
- Trace Apache/nginx/lighttpd rules, CGI mappings, environment variables, and shell/Perl/Python/PHP scripts.
- Search command construction beyond obvious shell metacharacters: arithmetic expansion, config files, response files, argument injection, newline/control characters, and third-party CLI parsing.
- Inspect support/debug functions, backup/restore, package install, log/telemetry processors, custom tags/templates, and native admin command runners.
- Compare configuration and init/upgrade changes alongside application code.
### Java/JVM and .NET
- Use [Vineflower](https://github.com/Vineflower/vineflower) for Java class/JAR reconstruction and `javap -c` to confirm ambiguous bytecode.
- Use official [ILSpy/ilspycmd](https://github.com/icsharpcode/ILSpy) for .NET assemblies and inspect IL/metadata when reconstructed C# is ambiguous.
- Do not build or run decompiler output, target assemblies/classes, bundled build scripts, or embedded resources in their associated target runtimes/viewers.
- Diff class/resource inventories before decompiled text to separate compiler/obfuscator noise from semantic changes.
### Native Binaries
- Use official [Ghidra](https://github.com/NationalSecurityAgency/ghidra) for strings/imports/xrefs/decompilation and reproducible headless projects.
- Use [BinDiff](https://github.com/google/bindiff) after manifest/package triage isolates the relevant native binaries, and keep the disassembler/BinExport version pair compatible across both sides.
- Confirm changed length, auth, command, parser, and file-handling conditions in assembly/runtime; decompiler types and similarity scores are hypotheses.
- Record architecture-specific calling convention, endian, alignment, libc, allocator, and mitigations.
Load `memory_corruption` for bounds/lifetime/disclosure findings and exploitability analysis. Load `protocol_reverse_engineering` for custom/stateful message formats.
## Version and Patch Analysis
Compare more than one adjacent pair when possible:
```text
older unaffected/unknown -> vulnerable -> first fixed -> current
```
- Build changed-file/package/config manifests first.
- Identify the security invariant introduced by the patch.
- Review every caller/sibling handler using the patched helper/parser.
- Check branch backports and inconsistent fixes across SKUs/architectures.
- Re-test the old structural condition on the fixed build and nearby routes.
- Inspect boot/runtime overlays and upgrade scripts if static diff shows no meaningful change.
- Distinguish one CVE from one code path; advisories may bundle several bugs or fix only the most exposed route.
Pair with `advisory_to_poc` for evidence classification, public-PoC decomposition, vulnerable/fixed controls, and detector handoff.
## Hardware, Virtualization, and Emulation
Record what the test environment omits:
- hardware security module/TPM/secure element and device-bound keys
- NIC/accelerator/driver behavior, DMA, endian/alignment, and kernel modules
- boot chain, secure boot, verified partitions, recovery mode, watchdog, and HA peer
- model-specific memory, allocator pressure, process limits, and service configuration
- virtual appliance differences from physical products
Full-system emulation can help recover routes and protocol behavior but often changes drivers, timing, entropy, memory layout, certificates, hardware identity, and mitigations. Treat emulation results as a separate platform and reproduce security-relevant behavior on the actual supported model when the claim depends on those properties.
Do not disable ASLR, canaries, signature checks, or other mitigations without labeling the resulting demonstration as lab-only and nonrepresentative of default exploitability.
## Physical-Lab Prerequisites
Have a recovery path before live-device work:
- console, serial, hypervisor, snapshot, or other known-good rollback method
- exact in-scope image/build and a way to reapply it
- isolated management network and controlled outbound connectivity
- process or watchdog visibility and a safe way to capture one request at a time
## Runtime Observation
Within an authorized lab, collect:
- process tree, executable/build ID, argv, cwd, users/groups/capabilities, open ports/sockets/files, mounts, namespaces/containers
- service logs, audit logs, core files, watchdog/restart events, and packet captures
- loaded mappings/libraries, relevant Unix sockets/file descriptors, and config source while sending one known request
- filesystem/process events while sending one known request
- boot/upgrade output and live configuration generated from templates/databases
Prefer observation that explains a static hypothesis. Do not install intrusive agents or attach a debugger to production equipment.
## Capability and Chain Mapping
Treat findings as product-context primitives:
- file read → configs, sessions, credentials, tokens, keys, topology
- SSRF/request → loopback APIs, sidecars, metadata, package agents
- file write → web roots, plugins, templates, restore packages, jobs, telemetry inputs
- auth bypass → support/admin command runners, package deployment, native operations
- parser disclosure → session/token/pointer material
- low-privilege identity → built-in management tools and trusted peer relationships
Inventory native product consumers before importing a generic exploit gadget. An appliance's normal backup, restore, diagnostic, package, scripting, or cluster function is frequently the shortest bridge between primitives.
## Deliverable
Include:
1. artifact provenance/hashes and complete SKU/version/platform/config matrix
2. extraction method and filesystem/boot/runtime architecture
3. listener/service/auth/trust-boundary map
4. changed-file/config/package manifest and relevant code path
5. external route/protocol through privileged operation and OS identity
6. hardware/emulation/mitigation constraints
7. vulnerable/fixed/negative-control behavior
8. adjacent handlers/branches/install states reviewed
9. tool versions, generated artifacts, and unresolved assumptions
## Common Errors
- Diffing different SKUs/architectures and attributing packaging noise to a security fix.
- Assuming extracted rootfs equals live state despite overlays, generation, or boot-time patches.
- Mapping only the web UI and missing auxiliary/custom/local listeners.
- Treating a hidden route as removed or a blocked route as a patched sink.
- Assuming fresh-install behavior covers upgraded systems with retained files/configuration.
- Calling a service pre-auth because a connection succeeds before a privileged operation is attempted.
- Treating emulator-only behavior or disabled mitigations as representative of a shipping device.
- Running an analyzed binary, extension, build script, or firmware helper on the analyst host.
## Summary
Appliances are integrated systems, not single applications. Preserve artifact lineage, extract safely, map boot/runtime state and every listener, trace edge configuration into code and privileged native features, compare fixes across branches and install states, and keep hardware/platform constraints attached to every finding.
@@ -0,0 +1,228 @@
---
name: memory-corruption
description: Native memory-safety analysis for stack and heap overflows, out-of-bounds access, uninitialized memory, use-after-free, integer and signedness errors, format strings, crash triage, exploitability constraints, and controlled lab validation
---
# Memory Corruption
Use this skill for authorized analysis of native parsers, network services, firmware daemons, libraries, and mixed web/native components where attacker-controlled bytes may violate memory safety.
Separate three questions throughout the work:
1. **Bug existence:** does an input cause an invalid read, write, lifetime violation, or disclosure?
2. **Primitive quality:** what bytes, address, length, timing, or object state can the attacker control or observe?
3. **Exploitability:** can that primitive bypass the target architecture, mitigations, allocator, protocol, and restart constraints?
A crash, connection close, watchdog restart, or sanitizer report proves neither instruction-pointer control nor RCE.
## Lab Boundary
Malformed-input and crash work is denial-of-service testing. Run it only against an explicitly authorized, restartable lab target with console/process visibility, health checks, rate ceilings, and a recovery procedure. Do not fuzz production services or automatically replay crash cases.
Analyze hostile binaries, cores, packet captures, and corpora inside an isolated environment. Do not execute an unknown sample merely because a debugger or decompiler imported it.
## Vulnerability Classes
### Bounds and Length Errors
- fixed destination with attacker-controlled copy/format length
- allocation based on one length and copy based on another
- off-by-one termination or delimiter handling
- nested length fields and cumulative-size overflow
- stack/heap out-of-bounds read or write
- negative length converted to unsigned, truncation between integer widths, or multiplication/addition overflow
- encoded/decoded/compressed size disagreement
### Initialization and Termination
- uninitialized stack/heap data returned in a response
- reused object/buffer retaining data from another request or tenant
- missing NUL termination followed by string length/format operations
- partial structure initialization with stale flags, pointers, or lengths
- padding, union, or serialization bytes copied beyond initialized fields
### Lifetime and Object Confusion
- use-after-free, double free, stale callback, iterator invalidation
- type/object confusion after parsing, casting, or virtual dispatch
- reference-count races and cross-thread ownership errors
- reallocation invalidating stored pointers
- constructor/destructor/finalizer behavior reached in an unexpected state
### Format and Variadic Errors
- attacker-controlled format string
- type/width mismatch in variadic arguments
- destination-size assumptions around `sprintf`-family calls
- logging/error paths that process attacker bytes after a partial parse
## Build the Input-to-Memory Model
Record:
```text
transport field -> parser type/width -> normalized value -> allocation
-> copy/read/format operation -> object/buffer -> later use
```
For each relevant field, capture:
- wire offset/path, endian, encoding, signedness, and declared versus actual size
- validation order and parser state required to reach the operation
- allocation expression and destination capacity
- copy/read/write expression and implicit casts
- terminator/padding/alignment behavior
- attacker-controlled byte alphabet and precision
- thread, connection, session, heap, and restart lifetime
Trace both source-to-sink and sink-to-source. Start from changed bounds checks or crash instructions when available, but reconstruct the minimum valid protocol state that reaches them.
## Source-Available Workflow
### Compiler Instrumentation
Build a lab-only target or minimal harness with the compiler's maintained sanitizers when source permits:
```bash
clang -g -O1 -fno-omit-frame-pointer \
-fsanitize=address,undefined \
harness.c parser.c -o parser-harness
```
- Keep the harness local and networkless; call the narrow parser/API directly.
- Preserve the exact compiler, flags, architecture, allocator, and dependencies.
- AddressSanitizer changes layout and timing. Reproduce important behavior on a representative unsanitized build under a debugger before drawing exploitability conclusions.
- UndefinedBehaviorSanitizer may report conditions that do not produce the deployed security impact; trace each report to attacker control and later use. It does not replace explicit arithmetic and cast review.
- For ordinary uninitialized-value hypotheses, use a separate MemorySanitizer build such as `-fsanitize=memory -fsanitize-memory-track-origins=2`; it requires an instrumented dependency set and is not interchangeable with ASan.
- For race-dependent ownership or refcount paths, use a separate ThreadSanitizer build only when concurrency is in scope; do not imply the sanitizer families compose cleanly into one representative build.
- Add regression cases for the minimized triggering input and neighboring non-triggering controls.
### Static Review
Search around input parsing for:
- `memcpy`, `memmove`, `strcpy`, `strcat`, `sprintf`, `snprintf`, `scanf` families
- manual cursor/end-pointer arithmetic and nested TLV/XML/string parsers
- `malloc/calloc/realloc/new` size arithmetic
- signed/unsigned conversions and narrowing casts
- length values stored in smaller fields or reused across decoded representations
- error cleanup, ownership transfer, callbacks, and asynchronous lifetime
- custom allocators, pools, slabs, ring buffers, and request-buffer reuse
Do not report a dangerous function name without proving attacker control, reachable state, capacity mismatch, and the actual deployed implementation.
## Binary-Only Workflow
1. Identify architecture, endian, ABI, OS/libc, compiler clues, and stripped/symbol state.
2. Record NX/DEP, ASLR/PIE, stack canaries, RELRO, CFI/PAC/CET, allocator hardening, seccomp/sandbox, privilege, and restart behavior.
3. Anchor on imports, strings, message IDs, error paths, new checks, crash PC, or advisory-relevant constants.
4. Trace length/copy/allocation dataflow in decompiler and assembly.
5. Record the deployed binary identity: build ID or hash, interpreter or loader, loaded modules/base addresses, allocator, and whether the runtime executable came from base image, overlay, bind mount, or update staging.
6. Reproduce under a debugger or emulator only when its environment matches the relevant parser and allocator behavior.
7. Compare vulnerable and fixed functions; describe the restored invariant and inspect sibling callers.
Use official [Ghidra](https://github.com/NationalSecurityAgency/ghidra) for cross-architecture static analysis and [BinDiff](https://github.com/google/bindiff) for function-level version comparison after package/file diffs narrow the target. Similarity scores and decompiled C are triage aids, not proof; confirm critical conditions in assembly and runtime evidence.
## Crash and Disclosure Triage
Preserve one known-good transcript and then minimize while keeping the framing, checksums, parser state, and negotiation required to reach the vulnerable operation. Identify the first invalid access, not only the eventual crash site. Use a distinctive non-executable pattern to measure overwrite offset or disclosure position, classify whether the observed effect is read, write, non-control-data, pointer/object, or control-state influence, and then repeat the same case on a representative unsanitized build plus fixed and negative controls.
For each case, record:
- exact minimized input and protocol transcript
- deterministic frequency and required heap/session preparation
- signal/exception, PC, faulting instruction bytes/disassembly, fault address, access type/size, registers, stack, loaded mappings/build IDs, and relevant object memory
- process versus worker crash, watchdog/restart, and external symptom
- corrupted object provenance and last known-valid parser state
- vulnerable/fixed/unaffected build behavior
- whether the same case under debugger/sanitizer changes outcome
Deduplicate by root cause, not only crash address. One overwrite may crash at many later consumers; one parser family may contain multiple distinct missing checks.
For disclosures, classify the returned bytes:
- predictable padding or constant data
- same-request content
- cross-request/tenant secrets
- heap/stack pointers useful against ASLR
- session tokens, keys, credentials, or application data
Derive detectors from response structure or a constant non-secret marker rather than collecting sensitive memory.
## Primitive Analysis
### Write Primitive
- location: fixed, relative, attacker-derived, heap-neighbor, object field, return/control data
- width and count: single byte/bit, bounded span, arbitrary length, repeated writes
- value control: exact, restricted alphabet, additive, terminator, pointer-derived
- timing/state: before validation, after free, race-dependent, heap-shape-dependent
- repeatability under default allocator and mitigations
### Read/Leak Primitive
- offset and length control
- termination rules and response encoding
- ability to repeat/advance across memory
- cross-request process reuse
- pointer or secret classification
- noise, truncation, and crash threshold
### Control-Flow/Object Primitive
- overwritten callback, vtable, length, non-control-data flag, pointer, credential/session reference, allocator metadata, saved return state, or interpreter structure
- required heap grooming/object placement
- available modules/gadgets and address disclosure
- thread/process privilege and sandbox boundary after control
- whether the attacker can only corrupt a field, or can also choose the dereference target and value later consumed
Document what remains constrained. “Arbitrary write” should not be used for a relative, partial, alphabet-limited, or race-only overwrite.
## Exploitability Matrix
| Dimension | Record |
|---|---|
| Reachability | listener, authentication, feature/config, valid prior state |
| Platform | architecture, endian, ABI, firmware model/SKU |
| Input | transport, maximum size, forbidden bytes, encoding/transforms |
| Primitive | read/write/control precision, repeatability, heap dependence |
| Mitigations | ASLR/PIE, NX, canary, RELRO, CFI/PAC/CET, allocator, sandbox |
| Process | privilege, chroot/container, worker isolation, watchdog/restart |
| Information | version fingerprint, pointer/module/heap leak availability |
| Reliability | attempts, races, connection/session persistence, crash side effects |
Rate exploitability separately from bug severity. A strong memory disclosure can enable a later control-flow bug; a large overflow may remain crash-only under the deployed constraints.
## Protocol and Patch Pairing
- Load `protocol_reverse_engineering` when valid negotiation/state is required before the vulnerable field.
- Load `advisory_to_poc` for vulnerable/fixed artifact matrices and patch-invariant review.
- Load `appliance_firmware` for rootfs, listener, runtime overlay, architecture, and device lifecycle mapping.
- Model transformation boundaries explicitly when the memory length or type changes across transport, parser, decoder, or native FFI layers.
## Validation Deliverable
Include:
1. exact vulnerable/fixed build, platform, configuration, and artifact hashes
2. minimized input plus complete protocol/parser prerequisites
3. source, IR/bytecode, or assembly trace from attacker field to invalid access, with the exact crashing process/build identity
4. debugger/sanitizer/core evidence and non-triggering control
5. primitive precision and constraints
6. mitigation, architecture, allocator, process, and restart analysis
7. bug-existence and exploitability conclusions stated separately
8. adjacent callers/parser family reviewed
## False Positives
- Connection close caused by protocol rejection, idle timeout, rate limit, or load balancer behavior.
- Process restart inferred from one failed request without process/console evidence.
- Sanitizer finding unreachable in the deployed feature, route, architecture, or configuration.
- Out-of-bounds read that returns only deterministic in-buffer padding, described as sensitive disclosure.
- Crash-only overwrite called RCE without a controlled data/control primitive and mitigation analysis.
- Decompiler type or buffer size accepted as ground truth without assembly/runtime confirmation.
- Lab build with mitigations disabled presented as representative of production.
## Summary
Memory-corruption research is constraint analysis. Trace exact bytes through length, allocation, copy, object lifetime, and later use; establish the read/write/control primitive; then evaluate architecture, mitigations, allocator, protocol, and process context independently from the mere existence of a crash.
-16
View File
@@ -112,19 +112,3 @@ def test_wait_for_agents_is_available_in_both_modes() -> None:
for interactive in (True, False): for interactive in (True, False):
agent = factory.build_strix_agent(is_root=True, interactive=interactive) agent = factory.build_strix_agent(is_root=True, interactive=interactive)
assert "wait_for_agents" in [t.name for t in agent.tools] assert "wait_for_agents" in [t.name for t in agent.tools]
def test_strict_tool_schemas_can_be_disabled_per_route() -> None:
"""Claude routes cap strict tools; the toolset must be sendable without strict."""
agent = factory.build_strix_agent(is_root=True, strict_tool_schemas=False)
function_tools = [t for t in agent.tools if isinstance(t, FunctionTool)]
assert function_tools
assert not any(t.strict_json_schema for t in function_tools)
def test_disabling_strict_leaves_shared_tools_untouched() -> None:
factory.build_strix_agent(is_root=True, strict_tool_schemas=False)
agent = factory.build_strix_agent(is_root=True)
assert any(t.strict_json_schema for t in agent.tools if isinstance(t, FunctionTool))
-15
View File
@@ -227,18 +227,3 @@ def test_resume_still_requires_targets_or_a_workspace(
cli_main.parse_arguments() cli_main.parse_arguments()
assert "has no targets_info" in capsys.readouterr().err assert "has no targets_info" in capsys.readouterr().err
def test_resume_non_object_run_json_exits(tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]) -> None:
monkeypatch.chdir(tmp_path)
run_dir = tmp_path / "strix_runs" / "pentest_abcd"
run_dir.mkdir(parents=True)
(run_dir / "run.json").write_text("[]", encoding="utf-8")
monkeypatch.setattr(sys, "argv", ["strix", "--resume", "pentest_abcd"])
with pytest.raises(SystemExit) as exc_info:
cli_main.parse_arguments()
assert exc_info.value.code == 2
captured = capsys.readouterr()
assert "run.json unreadable" in captured.err
assert "not an object" in captured.err
-22
View File
@@ -9,7 +9,6 @@ from strix.config.models import (
RECOMMENDED_MODEL_NAMES, RECOMMENDED_MODEL_NAMES,
is_recommended_or_frontier_model, is_recommended_or_frontier_model,
request_timeout_extra_args, request_timeout_extra_args,
supports_strict_tool_schemas,
) )
@@ -91,24 +90,3 @@ def test_frontier_model_families_are_accepted(model_name: str) -> None:
) )
def test_non_frontier_models_are_rejected(model_name: str) -> None: def test_non_frontier_models_are_rejected(model_name: str) -> None:
assert not is_recommended_or_frontier_model(model_name) assert not is_recommended_or_frontier_model(model_name)
@pytest.mark.parametrize(
"model_name",
[
"anthropic/claude-sonnet-4-6",
"bedrock/anthropic.claude-opus-4-8-v1:0",
"vertex_ai/claude-sonnet-5",
"Sonnet-5",
],
)
def test_claude_routes_reject_strict_tool_schemas(model_name: str) -> None:
assert not supports_strict_tool_schemas(model_name)
@pytest.mark.parametrize(
"model_name",
["openai/gpt-5.4", "gpt-5.4", "gemini/gemini-3.1-pro-preview", "deepseek/deepseek-v4"],
)
def test_other_routes_keep_strict_tool_schemas(model_name: str) -> None:
assert supports_strict_tool_schemas(model_name)
+2 -26
View File
@@ -13,7 +13,7 @@ from agents.tool import ToolOutputImage
from strix.config.settings import DEFAULT_MAX_TURNS from strix.config.settings import DEFAULT_MAX_TURNS
from strix.interface.tui.backend.controller import TuiController from strix.interface.tui.backend.controller import TuiController
from strix.interface.tui.backend.projection import bounded_state_projection, terminal_projection from strix.interface.tui.backend.projection import terminal_projection
from strix.interface.tui.backend.protocol import ( from strix.interface.tui.backend.protocol import (
MAX_COMMAND_BYTES, MAX_COMMAND_BYTES,
PROTOCOL_CAPABILITIES, PROTOCOL_CAPABILITIES,
@@ -215,11 +215,7 @@ def test_unicode_heavy_setup_state_stays_within_control_frame_limit() -> None:
"Any", "Any",
SimpleNamespace( SimpleNamespace(
caido_url="https://例え.example/" + "" * 10_000, caido_url="https://例え.example/" + "" * 10_000,
get_total_llm_usage=lambda: { get_total_llm_usage=lambda: {f"model-{index}": "" * 10_000 for index in range(20)},
"total_tokens": 720_400,
"cost": 20.0,
**{f"model-{index}": "🔒" * 10_000 for index in range(20)},
},
), ),
) )
server = TuiBackendServer(controller) server = TuiBackendServer(controller)
@@ -230,26 +226,6 @@ def test_unicode_heavy_setup_state_stays_within_control_frame_limit() -> None:
assert len(encoded) <= MAX_COMMAND_BYTES assert len(encoded) <= MAX_COMMAND_BYTES
assert "🔒".encode() in encoded assert "🔒".encode() in encoded
assert snapshot["projection_truncated"] is True assert snapshot["projection_truncated"] is True
assert snapshot["usage"] == {"total_tokens": 720_400, "cost": 20.0}
def test_defensive_state_projection_preserves_usage_summary() -> None:
controller = TuiController(args())
controller.report_state = cast(
"Any",
SimpleNamespace(
caido_url=None,
get_total_llm_usage=lambda: {"total_tokens": 720_400, "cost": 20.0},
),
)
state = controller.snapshot()
state["provider"] = None
state["future_oversized_field"] = "x" * 100_000
snapshot = bounded_state_projection(state)
assert snapshot["projection_truncated"] is True
assert snapshot["usage"] == {"total_tokens": 720_400, "cost": 20.0}
@pytest.mark.asyncio @pytest.mark.asyncio
+7 -49
View File
@@ -11,7 +11,6 @@ from typing import TYPE_CHECKING
from urllib.parse import urlsplit from urllib.parse import urlsplit
from strix.core.paths import latest_run_dir, runs_base_dir from strix.core.paths import latest_run_dir, runs_base_dir
from strix.interface.viewer.cli import run_view
from strix.interface.viewer.server import serve from strix.interface.viewer.server import serve
from strix.interface.viewer.transcript import ( from strix.interface.viewer.transcript import (
build_run_state, build_run_state,
@@ -49,31 +48,6 @@ def test_latest_run_dir_none_when_no_runs(tmp_path: Path, monkeypatch: pytest.Mo
assert runs_base_dir() == tmp_path / "strix_runs" assert runs_base_dir() == tmp_path / "strix_runs"
def test_view_cli_help_includes_host(capsys: pytest.CaptureFixture[str]) -> None:
try:
run_view(["--help"])
except SystemExit as exc:
assert exc.code == 0
else:
raise AssertionError("--help should exit")
help_text = capsys.readouterr().out
assert "--host HOST" in help_text
assert "0.0.0.0" in help_text
def test_server_can_bind_all_ipv4_interfaces(tmp_path: Path) -> None:
run_dir = _make_run(tmp_path, "remote", status="running", end_time=None)
httpd, url, _ = serve(run_dir, host="0.0.0.0", open_browser=False)
try:
assert httpd.server_address[0] == "0.0.0.0"
assert url == f"http://0.0.0.0:{httpd.server_address[1]}"
finally:
httpd.shutdown()
httpd.server_close()
def test_latest_run_dir_picks_newest_by_record_mtime( def test_latest_run_dir_picks_newest_by_record_mtime(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None: ) -> None:
@@ -199,15 +173,14 @@ def test_server_serves_api_and_static(tmp_path: Path, monkeypatch: pytest.Monkey
(assets / "assets" / "app.js").write_text("console.log(1)", encoding="utf-8") (assets / "assets" / "app.js").write_text("console.log(1)", encoding="utf-8")
monkeypatch.setattr("strix.interface.viewer.server.bundle_dir", lambda: assets) monkeypatch.setattr("strix.interface.viewer.server.bundle_dir", lambda: assets)
httpd, url, token = serve(run_dir, open_browser=False) httpd, url, _ = serve(run_dir, open_browser=False)
try: try:
cookie = _session_cookie(url, token) status, ctype, body = _get(f"{url}/api/run")
status, ctype, body = _get(f"{url}/api/run", cookie=cookie)
assert status == 200 assert status == 200
assert "application/json" in ctype assert "application/json" in ctype
assert json.loads(body)["finished"] is True assert json.loads(body)["finished"] is True
status, _, body = _get(f"{url}/api/transcript", cookie=cookie) status, _, body = _get(f"{url}/api/transcript")
assert {a["id"] for a in json.loads(body)["agents"]} == {"root", "child"} assert {a["id"] for a in json.loads(body)["agents"]} == {"root", "child"}
# Real asset is served. # Real asset is served.
@@ -456,22 +429,6 @@ def test_unauthorized_client_cannot_acquire_capability(
httpd.server_close() httpd.server_close()
def test_run_data_requires_session(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
run_dir = _make_run(tmp_path, "private", status="completed", end_time="2026-01-01T00:00:00Z")
_bundle(tmp_path, monkeypatch)
httpd, url, token = serve(run_dir, open_browser=False)
try:
cookie = _session_cookie(url, token)
for path in ("/api/run", "/api/vulnerabilities", "/api/report", "/api/transcript"):
assert _get_status(url + path) == 403, path
assert _get_status(url + path, cookie=f"{_cookie_name(url)}=wrong") == 403, path
assert _get_status(url + path, cookie=cookie) == 200, path
finally:
httpd.shutdown()
httpd.server_close()
def test_auth_status_reflects_expiry(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None: def test_auth_status_reflects_expiry(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
run_dir = _make_run(tmp_path, "status", status="running", end_time=None) run_dir = _make_run(tmp_path, "status", status="running", end_time=None)
_bundle(tmp_path, monkeypatch) _bundle(tmp_path, monkeypatch)
@@ -604,10 +561,11 @@ def test_historical_run_data_requires_verification(
httpd, url, token = serve(launched, open_browser=False) httpd, url, token = serve(launched, open_browser=False)
try: try:
# The launched run needs the session capability, but not email verification. # The launched run is always viewable, no verification and no cookie.
assert _get_status(f"{url}/api/run") == 403 status, _, _ = _get(f"{url}/api/run")
assert status == 200
cookie = _session_cookie(url, token) cookie = _session_cookie(url, token)
assert _get_status(f"{url}/api/run", cookie=cookie) == 200
# A different run needs the session capability first: a cookie-less # A different run needs the session capability first: a cookie-less
# caller is forbidden even once the machine is verified. # caller is forbidden even once the machine is verified.