mirror of
https://github.com/usestrix/strix.git
synced 2026-08-25 04:12:37 +02:00
Rewrite OSS docs for ASD-STE100 Simplified Technical English
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: "Azure OpenAI"
|
||||
description: "Configure Strix with OpenAI models via Azure"
|
||||
description: "Configure Strix with OpenAI models through Azure"
|
||||
---
|
||||
|
||||
## Setup
|
||||
@@ -19,7 +19,7 @@ export AZURE_API_VERSION="2025-11-01-preview"
|
||||
| `STRIX_LLM` | `azure/<your-deployment-name>` |
|
||||
| `AZURE_API_KEY` | Your Azure OpenAI API key |
|
||||
| `AZURE_API_BASE` | Your Azure OpenAI endpoint URL |
|
||||
| `AZURE_API_VERSION` | API version (e.g., `2025-11-01-preview`) |
|
||||
| `AZURE_API_VERSION` | API version, such as `2025-11-01-preview` |
|
||||
|
||||
## Example
|
||||
|
||||
@@ -33,5 +33,5 @@ export AZURE_API_VERSION="2025-11-01-preview"
|
||||
## Prerequisites
|
||||
|
||||
1. Create an Azure OpenAI resource
|
||||
2. Deploy a model (e.g., GPT-5.4)
|
||||
2. Deploy a model, such as GPT-5.4
|
||||
3. Get the endpoint URL and API key from the Azure portal
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: "AWS Bedrock"
|
||||
description: "Configure Strix with models via AWS Bedrock"
|
||||
description: "Configure Strix with models through AWS Bedrock"
|
||||
---
|
||||
|
||||
## Installation
|
||||
@@ -17,7 +17,7 @@ pipx install "strix-agent[bedrock]"
|
||||
export STRIX_LLM="bedrock/anthropic.claude-4-5-sonnet-20251022-v1:0"
|
||||
```
|
||||
|
||||
No API key required—uses AWS credentials from environment.
|
||||
Strix does not require an API key. Strix uses AWS credentials from the environment.
|
||||
|
||||
## Authentication
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@ description: "Run Strix with self-hosted LLMs for privacy and air-gapped testing
|
||||
|
||||
Running Strix with local models allows for completely offline, privacy-first security assessments. Data never leaves your machine, making this ideal for sensitive internal networks or air-gapped environments.
|
||||
|
||||
## Privacy vs Performance
|
||||
## Privacy Versus Performance
|
||||
|
||||
| Feature | Local Models | Cloud Models (GPT-5/Claude 4.5) |
|
||||
|---------|--------------|--------------------------------|
|
||||
@@ -14,11 +14,11 @@ Running Strix with local models allows for completely offline, privacy-first sec
|
||||
| **Reasoning** | Lower (struggles with agents) | State-of-the-art |
|
||||
| **Setup** | Complex (GPU required) | Instant |
|
||||
|
||||
<Warning>
|
||||
**Compatibility Note**: Strix relies on advanced agentic capabilities (tool use, multi-step planning, self-correction). Most local models, especially those under 70B parameters, struggle with these complex tasks.
|
||||
<Note>
|
||||
Strix requires advanced agent capabilities, including tool use, multi-step planning, and self-correction. Most local models under 70B parameters struggle with these tasks.
|
||||
|
||||
For critical assessments, we strongly recommend using state-of-the-art cloud models like **Claude 4.5 Sonnet** or **GPT-5**. Use local models only when privacy is the absolute priority.
|
||||
</Warning>
|
||||
Critical assessments often require capable cloud models. Local models suit assessments where privacy has priority.
|
||||
</Note>
|
||||
|
||||
## Ollama
|
||||
|
||||
@@ -39,9 +39,7 @@ For critical assessments, we strongly recommend using state-of-the-art cloud mod
|
||||
|
||||
### Recommended Models
|
||||
|
||||
We recommend these models for the best balance of reasoning and tool use:
|
||||
|
||||
**Recommended models:**
|
||||
We recommend these models for a good balance of reasoning and tool use:
|
||||
- **Qwen3 VL** (`ollama pull qwen3-vl`)
|
||||
- **DeepSeek V3.1** (`ollama pull deepseek-v3.1`)
|
||||
- **Devstral 2** (`ollama pull devstral-2`)
|
||||
@@ -59,7 +57,7 @@ export LLM_API_BASE="http://localhost:1234/v1" # Adjust port as needed
|
||||
|
||||
Some OpenAI-compatible gateways require extra HTTP headers (for attribution or
|
||||
tenant routing) alongside the bearer token. Set them with `LLM_EXTRA_HEADERS` as
|
||||
a JSON object — they are sent on every request:
|
||||
a JSON object. Strix sends these headers on every request:
|
||||
|
||||
```bash
|
||||
export STRIX_LLM="openai/your-model"
|
||||
@@ -69,12 +67,12 @@ export LLM_EXTRA_HEADERS='{"X-Feature-Key":"value","X-Tenant":"acme"}'
|
||||
```
|
||||
|
||||
For endpoints behind a private CA, point Strix at your certificate bundle with
|
||||
the standard `SSL_CERT_FILE=/path/to/ca-bundle.pem` — never disable TLS
|
||||
the standard `SSL_CERT_FILE=/path/to/ca-bundle.pem`. Do not disable TLS
|
||||
verification against a real endpoint.
|
||||
|
||||
## Tool calling must return structured `tool_calls`
|
||||
|
||||
Strix is entirely tool-driven: every working turn must be a **native** function/tool call. If your inference server returns the tool call as plain assistant text instead of a structured `tool_calls` field, Strix never sees a call it can execute, so the agent makes no real progress — it re-prompts the model for a tool call and gives up once its recovery attempts are exhausted.
|
||||
Strix requires a **native** function or tool call during each working turn. If the inference server returns plain assistant text, Strix cannot execute the call. The agent then requests another tool call and stops after its recovery attempts end.
|
||||
|
||||
This is almost always an **inference-server configuration** problem, not a model or Strix problem. Common symptoms are the model printing a call as text such as:
|
||||
|
||||
@@ -84,19 +82,19 @@ exec_command(cmd="nmap ...", timeout=180)
|
||||
{"action": "exec_command", "params": {"cmd": "nmap ..."}}
|
||||
```
|
||||
|
||||
The fix belongs on the inference server: it must be configured to parse the model's tool tokens into structured `tool_calls`. A correctly configured endpoint either returns a structured call or rejects the request outright — it never leaks the call as text.
|
||||
Configure the inference server to parse tool tokens into structured `tool_calls`. A configured endpoint returns a structured call or rejects the request. It does not return the call as text.
|
||||
|
||||
### Fixes by server
|
||||
|
||||
**llama.cpp (`llama-server`)**
|
||||
- Run with `--jinja` and a correct tool-use chat template (`--chat-template` / `--chat-template-file` matching the model). Recent builds enable `--jinja` by default — **upgrade** if yours doesn't.
|
||||
- For thinking models, align or disable reasoning (`--reasoning-format`, `-rea off`) so it doesn't break tool-call parsing.
|
||||
- A low temperature (e.g. `--temp 0.2`) improves tool-call reliability.
|
||||
- Run with `--jinja` and a tool-use chat template that matches the model. Recent builds enable `--jinja` by default. Upgrade if yours does not.
|
||||
- For thinking models, align or disable reasoning with `--reasoning-format` or `-rea off`.
|
||||
- Set a low temperature, such as `--temp 0.2`, to improve tool-call reliability.
|
||||
|
||||
**Ollama**
|
||||
- Use a recent Ollama and a model whose template wires tools. Modern Ollama refuses tools (`tools param requires --jinja flag`) if the template lacks tool support.
|
||||
- For reasoning models (e.g. qwen3), disable the model's **thinking** mode — thinking left on frequently pushes the tool call into the text `content` instead of the structured `tool_calls` field. Turn it off on the Ollama side (a non-thinking model variant, or `think: false` in the model's parameters / `Modelfile`).
|
||||
- Raise **`num_ctx`** to at least 16k–32k. Strix sends a large system prompt plus many tool schemas; at Ollama's small default context the tool definitions are truncated out of the prompt and the model stops emitting valid calls. A short test prompt can look fine while a real scan fails, so set this explicitly rather than inferring it from a quick check.
|
||||
- Use a recent Ollama version and a model whose template supports tools. Ollama refuses tools when the template lacks tool support.
|
||||
- For reasoning models such as qwen3, disable **thinking** mode. Thinking mode can move tool calls into `content` instead of `tool_calls`. Disable it in Ollama with a non-thinking model or `think: false`.
|
||||
- Raise **`num_ctx`** to at least 16k. Strix sends a large system prompt and many tool schemas. A small context can truncate the tool definitions and stop valid calls.
|
||||
|
||||
**vLLM**
|
||||
- Start with `--enable-auto-tool-choice`, a matching `--tool-call-parser` (`hermes`, `qwen3_xml`, or `llama3_json`), and a matching `--reasoning-parser` for reasoning models.
|
||||
|
||||
@@ -3,7 +3,7 @@ title: "Novita AI"
|
||||
description: "Configure Strix with Novita AI models"
|
||||
---
|
||||
|
||||
[Novita AI](https://novita.ai) provides fast, cost-efficient inference for open-source models via an OpenAI-compatible API.
|
||||
[Novita AI](https://novita.ai) provides fast, cost-efficient inference for open-source models through an OpenAI-compatible API.
|
||||
|
||||
## Setup
|
||||
|
||||
@@ -29,7 +29,7 @@ export LLM_API_BASE="https://api.novita.ai/openai"
|
||||
|
||||
## Benefits
|
||||
|
||||
- **Cost-efficient** — Competitive pricing with per-token billing
|
||||
- **OpenAI-compatible** — Drop-in replacement using `LLM_API_BASE`
|
||||
- **Large context** — Models support up to 262k token context windows
|
||||
- **Function calling** — All listed models support tool/function calling
|
||||
- **Cost-efficient:** Competitive pricing with per-token billing
|
||||
- **OpenAI-compatible:** Drop-in replacement using `LLM_API_BASE`
|
||||
- **Large context:** Models support up to 262k token context windows
|
||||
- **Function calling:** All listed models support tool/function calling
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: "OpenRouter"
|
||||
description: "Configure Strix with models via OpenRouter"
|
||||
description: "Configure Strix with models through OpenRouter"
|
||||
---
|
||||
|
||||
[OpenRouter](https://openrouter.ai) provides access to 100+ models from multiple providers through a single API.
|
||||
@@ -31,7 +31,7 @@ Access any model on OpenRouter using the format `openrouter/<provider>/<model>`:
|
||||
|
||||
## Benefits
|
||||
|
||||
- **Single API** — Access models from OpenAI, Anthropic, Google, Meta, and more
|
||||
- **Fallback routing** — Automatic failover between providers
|
||||
- **Cost tracking** — Monitor usage across all models
|
||||
- **Higher rate limits** — OpenRouter handles provider limits for you
|
||||
- **Single API:** Access models from OpenAI, Anthropic, Google, Meta, and more
|
||||
- **Fallback routing:** Automatic failover between providers
|
||||
- **Cost tracking:** Monitor usage across all models
|
||||
- **Higher rate limits:** OpenRouter handles provider limits for you
|
||||
|
||||
@@ -44,13 +44,13 @@ See the [Local Models guide](/llm-providers/local) for setup instructions and re
|
||||
Access 100+ models through a single API.
|
||||
</Card>
|
||||
<Card title="Google Vertex AI" href="/llm-providers/vertex">
|
||||
Gemini 3 models via Google Cloud.
|
||||
Gemini 3 models through Google Cloud.
|
||||
</Card>
|
||||
<Card title="AWS Bedrock" href="/llm-providers/bedrock">
|
||||
Claude and Titan models via AWS.
|
||||
Claude and Titan models through AWS.
|
||||
</Card>
|
||||
<Card title="Azure OpenAI" href="/llm-providers/azure">
|
||||
GPT-5.4 via Azure.
|
||||
GPT-5.4 through Azure.
|
||||
</Card>
|
||||
<Card title="Local Models" href="/llm-providers/local">
|
||||
Llama 4, Mistral, and self-hosted models.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: "Google Vertex AI"
|
||||
description: "Configure Strix with Gemini models via Google Cloud"
|
||||
description: "Configure Strix with Gemini models through Google Cloud"
|
||||
---
|
||||
|
||||
## Installation
|
||||
@@ -17,7 +17,7 @@ pipx install "strix-agent[vertex]"
|
||||
export STRIX_LLM="vertex_ai/gemini-3-pro-preview"
|
||||
```
|
||||
|
||||
No API key required—uses Google Cloud Application Default Credentials.
|
||||
Strix does not require an API key. Strix uses Google Cloud Application Default Credentials.
|
||||
|
||||
## Authentication
|
||||
|
||||
|
||||
Reference in New Issue
Block a user