23 KiB
Design: Matty — forkless' annoying Mattermost sidekick
Status: v0.1.17 — Implemented and deployed
Last updated: 2026-06-17
Bot username: @matty
Default model: dolphin-llama3:latest (configurable in System Console)
1. Purpose & Scope
1.1 What It Is
Matty is a Mattermost (Community Edition) plugin that lets users chat
with remote LLMs through an Ollama backend. Users address the bot by
@-mentioning it in any channel (e.g. @matty write a poem). The
bot responds only when addressed and disengages naturally — it never
reads every message in the room. Once addressed, the user enters an
active session where follow-up messages continue the conversation
without needing to re-@mention. Sessions end on stop words ("thanks",
"bye"), explicit @matty stop, or 5 minutes of inactivity.
A slash command (/ai) remains available as an alternative entry point.
1.2 What It Is Not (Explicit Out-of-Scope)
| Out of scope | Rationale |
|---|---|
| Hosting its own LLM inference | Ollama is the sole backend; the plugin is a thin proxy |
| Multi-backend abstraction (OpenAI, Anthropic, etc.) | If needed later, abstract behind an LLMProvider interface — not now |
| Training / fine-tuning | Ollama handles that |
| Channel whisper mode (replying to every message) | The agent only activates on @-mention or active session. It never reads unprompted messages. |
| Bot account management | Plugin creates the bot account automatically via API.CreateBot() |
| Webapp UI | No frontend — @-mention + /ai covers all interaction |
1.3 Target Audience
- Teams self-hosting Mattermost who want AI assistant access without data leaving their infrastructure
- Users comfortable with @-mentioning a bot the same way they @-mention a human teammate
2. Architecture Overview
2.1 Component Diagram (text)
┌──────────────────────────────────────────────────────────────┐
│ Mattermost Server (Community Edition) │
│ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ Matty Plugin │ │
│ │ │ │
│ │ ┌──────────────────┐ ┌────────────────────────┐ │ │
│ │ │ MessageHasBeen │──▶│ Plugin (Go) │ │ │
│ │ │ Posted (primary) │ │ │ │ │
│ │ └──────────────────┘ │ ┌──────────────────┐ │ │ │
│ │ ┌──────────────────┐ │ │ OllamaClient │──┼───┼───┼──▶ Ollama
│ │ │ SlashCommand │──▶│ └──────────────────┘ │ │ │ (remote)
│ │ │ /ai (alt.) │ │ ┌──────────────────┐ │ │ │
│ │ └──────────────────┘ │ │ ConversationStore│ │ │ │
│ │ ┌──────────────────┐ │ └──────────────────┘ │ │ │
│ │ │ EngagementEngine │ │ ┌──────────────────┐ │ │ │
│ │ │ sleep / active │──▶│ │ RateLimiter │ │ │ │
│ │ └──────────────────┘ │ └──────────────────┘ │ │ │
│ │ ┌──────────────────┐ │ ┌──────────────────┐ │ │ │
│ │ │ PersonaLibrary │ │ │ Config (KV) │ │ │ │
│ │ │ resolvePersona() │ │ └──────────────────┘ │ │ │
│ │ └──────────────────┘ │ │ │ │
│ │ ┌──────────────────┐ │ │ │ │
│ │ │ Bot Account │ │ │ │ │
│ │ │ (auto-created) │ │ │ │ │
│ │ └──────────────────┘ │ │ │ │
│ │ ┌──────────────────┐ │ │ │ │
│ │ │ Webapp (none) │ │ │ │ │
│ │ └──────────────────┘ │ │ │ │
│ └──────────────────────────────────────────────────────┘ │
│ │ │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ KV Store (per-plugin) │ │
│ │ • ctx_{user}_{channel} — conversation history │ │
│ │ • persona_{user}_{ch} — persona override │ │
│ └──────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────────┘
2.2 Request Flow (happy path — @-mention)
User types "@matty write a release note for v2.1"
│
▼
MessageHasBeenPosted fires
│
▼
EngagementEngine.Wake(userId, channelId) → mark session Active
│
▼
@-mention detected → strip "@matty " prefix → "write a release note for v2.1"
│
▼
RateLimiter.Allow(userId) → 200
│
▼
Check stop words → no match
│
▼
Check persona override (KV store) → resolvePersona() → full description
│
▼
ConversationStore.BuildMessages(userId, channelId, prompt) → up to 10 turns
│
▼
Trim context based on Brevity (e.g. brevity=1 → keep only current message)
│
▼
Prepend system prompt: "You are {persona}. Respond in character. ... Reply in {N} sentences."
│
▼
Calculate token cap: maxTokens = auto from brevity, or manual override
│
▼
OllamaClient.ChatCompletionStream(model, messages, options)
│
▼
Ollama returns ndjson stream → accumulate all tokens
│
▼
Create single post with full response (not a thread reply, not live-updated)
2.3 Engagement Session Flow (follow-up without @-mention)
User: "@matty hello" → bot responds, session starts
User: "tell me a joke" (no @) → IsActive() → true → bot responds
User: "what about weather?" (no @) → IsActive() → true → bot responds
User: "thanks" → stopWordMatch → Sleep() → bot signs off
User: "hello again" (no @, 10 min later)→ IsActive() → false → ignored
User: "@matty hello again" → Wake() → bot responds, new session
2.4 Error Flow
| Failure point | Behaviour |
|---|---|
| Ollama unreachable | Post: "Matty is asleep. @matty when you need me." |
| Rate limit hit | Post: "Rate limit exceeded. Please wait before sending another request." |
| Invalid persona | Ephemeral: "Unknown persona. Available: ..." |
| Timeout (>30 s) | Post truncated with "(Matty is asleep...)" |
3. Plugin Manifest & Identity
3.1 plugin.json
Current settings (v0.1.17):
| Key | Type | Default | Purpose |
|---|---|---|---|
OllamaURL |
text | http://localhost:11434 |
Ollama server endpoint |
DefaultModel |
dropdown | dolphin-llama3:latest |
Model for chat |
Persona |
dropdown | assistant |
Character for the bot |
RepeatPenalty |
number | 11 (1.1) |
Penalise repeated tokens |
FrequencyPenalty |
number | 1 (0.1) |
Penalise frequent phrases |
Brevity |
dropdown | 3 |
Max sentences |
DebugLogging |
bool | false |
Enable verbose logging |
CustomPersona |
text | "" |
Override dropdown persona |
MaxTokens |
number | 0 |
Manual token cap (0 = auto from Brevity) |
RateLimitPerMinute |
number | 0 |
Requests/min/user (0 = unlimited) |
AllowedUserIDs |
text | "" |
Restrict to specific users |
StopWords |
text | thanks, bye, goodbye, done, stop, quit, end |
Disengagement triggers |
3.2 Versioning
Follow release-workflow skill: v0.1.x for small iterations. Bump PATCH for
bug fixes, MINOR for features. No MAJOR bump until stable release.
4. Server-Side Components
4.1 Plugin Lifecycle Hooks
| Hook | Purpose |
|---|---|
OnActivate() |
Register /ai command, create bot account via ensureBot(), initialise OllamaClient, ConversationStore, RateLimiter, EngagementEngine |
OnConfigurationChange() |
Re-read System Console settings, validate OllamaURL, sync debug logging flag to OllamaClient |
MessageHasBeenPosted() |
Primary entry point — detect @-mentions and active sessions, delegate to EngagementEngine |
ExecuteCommand() |
Alternative entry point — intercept /ai and subcommands (/ai persona, /ai new, /ai help) |
Implemented() |
Declares handled hooks: OnActivate, OnDeactivate, OnConfigurationChange, ExecuteCommand, MessageHasBeenPosted |
4.2 Interaction Model
4.2.1 @-mention (primary)
User types @matty <prompt> in any channel. The plugin:
- Detects the @-mention via
MessageHasBeenPosted→ string match on@matty - Strips the bot username from the prompt
- Checks stop words, rate limit, authorization
- Looks up persona override from KV store (set via
/ai persona <name>) - Builds conversation context from
ConversationStore(up to 10 turns) - Trims context based on Brevity setting (e.g. brevity 1 = only current message)
- Prepends persona system prompt with general instructions + brevity
- Calculates token cap (auto from Brevity or manual MaxTokens override)
- Wakes the
EngagementEnginefor this user+channel - Sends to Ollama via
ChatCompletionStream - Accumulates all tokens, creates a single channel post with the full response
When the bot does NOT respond:
- Messages without
@mattywhile the EngagementEngine is Sleeping - Messages during the Disengaging → Sleeping transition
- Messages from other bot accounts (infinite-loop guard)
4.2.2 Engagement Session (continuation)
Once a user is in an Active session (via @-mention or /ai), their subsequent
non-@-mention messages in the same channel are treated as continuations:
User enters Active → subsequent messages auto-continue → stop word → Sleep
The session expires after 5 minutes of inactivity.
4.2.3 Slash command (alternative)
/ai <prompt> works as a fallback. Behaviour is identical.
Full syntax:
/ai <prompt> — Chat with default model
/ai persona — Show current active persona
/ai persona list — List all available personas with descriptions
/ai persona <name> — Switch persona (persisted)
/ai new — Reset conversation context + end session
/ai help — Show help
4.3 OllamaClient
Located in server/ollama.go (package main, flat file).
Key methods:
| Method | Purpose |
|---|---|
ChatCompletion(req) |
Non-streaming call (not used by main flow) |
ChatCompletionStream(req) |
Primary — returns < -chan StreamEvent, reads ndjson |
ListModels() |
Returns available model names |
Ping() |
Health check |
The client accepts a logFunc callback for debug logging (hooked to
p.API.LogInfo when DebugLogging is enabled).
4.4 ConversationStore
Stored in KV store per ctx_{userID}_{channelID}. Contains the last 10
messages (5 user + 5 assistant). Context is trimmed further based on
Brevity setting before sending to Ollama.
| Brevity | Messages kept |
|---|---|
| 1 | 1 (current only) |
| 2 | 2 (last exchange) |
| 3 | 4 (last 2 exchanges) |
| 5 | 6 |
| 10 | 10 |
4.5 RateLimiter
Token-bucket per user. Configurable from System Console. Default: 0 (unlimited). Resets on config change. Stored in memory.
4.6 EngagementEngine
Tracks active conversations per user+channel. Two states:
| State | Behaviour | Entered when |
|---|---|---|
| Sleeping | Ignores all messages | Plugin start, stop word, /ai new, 5 min timeout |
| Active | Responds to @-mentions and follow-ups | @-mention or /ai |
The timeout is 5 minutes of inactivity. On expiry, the session is silently removed — next @-mention starts fresh.
4.7 Persona System
Located in server/persona.go. A library of 19 personas mapped by stable
short IDs:
assistant a helpful assistant. Be concise, 2-3 sentences max...
canadian an overly polite Canadian who apologises excessively...
carlin George Carlin. Acerbic, observational, profane comedy...
dnd a Dungeons & Dragons dungeon master...
infomercial a 1980s TV infomercial host. Over-enthusiastic, loud...
jester a medieval court jester who mocks everyone...
mime a French mime. Respond with dramatic pauses...
motivational a motivational speaker who turns everything into a life lesson...
noir a noir detective who answers every question with another...
philosopher a philosopher who overcomplicates everything...
pirate a pirate captain. Use pirate slang...
pirro a drunk Jeanine Pirro. Slur your words...
ramsay Gordon Ramsay but for code reviews. Aggressive chef...
scammer a tech support scammer from 2008. Alarmist, pushy...
shakespeare William Shakespeare himself. Respond in iambic pentameter...
surfer a surfer dude explaining complex topics badly...
sysadmin a grumpy sysadmin who thinks every question is stupid...
tiredparent a tired parent who's given up on life...
trump Donald Trump. Boastful, hyperbolic, third-person...
Priority: KV override (/ai persona <id>) → CustomPersona (System Console text field) → Persona dropdown.
resolvePersona(id) looks up the short ID in the library. Returns empty
for unknown IDs (so /ai persona shakespear is properly rejected).
4.8 Debug Logging
When DebugLogging is enabled in System Console, the plugin logs:
| Log message | Source | What it contains |
|---|---|---|
"Config loaded" |
configuration.go |
All settings on save |
"Chat request" |
hooks.go |
13 fields: model, url, persona, brevity, tokens, repeat, freq, rate limit, allowed users, stop words, full system prompt |
"Ollama API call" |
ollama.go |
Full request body sent to Ollama (model, messages, options) |
5. Ollama Integration
5.1 API Mapping
| Plugin Action | Ollama Endpoint | Method | Stream |
|---|---|---|---|
| Chat completion | /api/chat |
POST | Yes |
| List models | /api/tags |
GET | No |
| Health check | /api/tags |
GET | No |
5.2 Streaming Strategy
- Send POST to
/api/chatwith"stream": true. - Read ndjson response body line by line.
- Accumulate all tokens in a buffer (no live post updates — avoids "Edited" spam).
- On
done:true, create a single channel post with the full content.
5.3 Token Auto-Calculation
Brevity auto-calculates the token cap. Formula: 100 + brevity × 75.
| Brevity | Auto tokens | Override with |
|---|---|---|
| 1 | 175 | MaxTokens > 175 |
| 2 | 250 | MaxTokens > 250 |
| 3 | 325 | MaxTokens > 325 |
| 5 | 475 | MaxTokens > 475 |
| 10 | 850 | MaxTokens > 850 |
| Unlimited (0) | Falls back to MaxTokens (0 = Ollama default) | MaxTokens > 0 |
5.4 Penalties
RepeatPenalty and FrequencyPenalty use whole numbers divided by 10:
| UI value | Sent to Ollama |
|---|---|
| 10 | 1.0 (none) |
| 11 | 1.1 (mild) |
| 12 | 1.2 (strong) |
| UI value | Sent to Ollama |
|---|---|
| 0 | 0.0 (none) |
| 1 | 0.1 (mild) |
| 5 | 0.5 (strong) |
6. Configuration
All configuration is done via System Console → Plugins → Matty. The 12 settings are described in §3.1.
OnConfigurationChange() reloads all settings when saved. Plugins do not
have a SavePluginConfiguration API in Mattermost v11.x, so /ai persona
uses the KV store to persist overrides instead.
7. Data Flow — Detailed
7.1 MessageHasBeenPosted (full lifecycle)
Step User / System Component
──── ────────────────────────────── ──────────────────
1 User types "@matty explain TCP"
2 MessageHasBeenPosted fires
3 Check: not bot's own post Loop guard
4 Check: not from OAuth bot Loop guard
5 Check: has @matty? yes Mention detection
6 Check: EngagementEngine.IsActive? EngagementEngine
7 Strip @matty → "explain TCP"
8 Check stop words StopWordSet
9 Check rate limit RateLimiter
10 Check authorization AllowedUserIDs
11 EngagementEngine.Wake() EngagementEngine
12 GetPersonaOverride → resolvePersona KV Store / personaLibrary
13 BuildMessages → conversation history ConversationStore
14 Trim context by Brevity hooks.go
15 Build system prompt + brevity hooks.go
16 Calculate token cap hooks.go
17 Build Ollama options (penalties) hooks.go
18 Log Chat request (if debug) hooks.go
19 ChatCompletionStream() OllamaClient
20 Accumulate tokens hooks.go
21 CreatePost with full response postReply()
8. Security Model
8.1 Input Sanitization
- Stop words checked before sending to Ollama (saves round-trip)
- Rate limiting per user (configurable, default: unlimited)
8.2 Network Security
OllamaURLvalidated as http/https at config load- TLS enabled by default
8.3 Authorization
AllowedUserIDsrestricts by Mattermost user ID- Empty = all authenticated users
- Bot account created by plugin, managed automatically
8.4 Infinite-Loop Guard
The plugin never processes messages from:
- Its own bot account (
post.UserId == botUserID) - Other OAuth bots (
post.IsFromOAuthBot())
9. Observability
9.1 Logging
When Debug Logging is enabled:
| Level | When |
|---|---|
INFO |
Chat request (all 13 settings), Config loaded, Ollama API call body |
ERROR |
Ollama unreachable, config validation failure, KV store error |
WARN |
Rate limit hit (with debug on) |
9.2 Error Reporting
- Ollama unreachable → "Matty is asleep. @matty when you need me."
- Rate limit → "Rate limit exceeded..."
- Unknown persona → "Unknown persona. Available: ..."
- Stream error → appended as "(Matty is asleep...)" in the response
10. Resolved & Open Design Decisions
Resolved
| # | Question | Decision |
|---|---|---|
| D1 | Reply style | Direct channel posts (not ephemeral, not threads) |
| D4 | Webapp UI | None — @-mention + /ai + /ai persona covers everything |
| D6 | Thread replies | EngagementEngine — active sessions track follow-ups without threads |
| D11 | Engagement scope | Per-user+channel — separate sessions per channel |
| D12 | Thread timeout | 5 minutes (not configurable yet) |
Still Open
| # | Question | Proposed |
|---|---|---|
| D2 | Context scope | Per-channel per-user (implemented) |
| D3 | Context turns | Currently 10 messages; Brevity trims further |
| D5 | Model governance | All models available via System Console dropdown |
| D7 | num_ctx per request |
Not implemented; Brevity controls token cap |
| D8 | Auto-pull models | No (returns error with list) |
| D9 | Prompt injection guard | Deferred — document risk |
| D10 | Metrics endpoint | Deferred |
11. Deployment & Lifecycle
11.1 Build
make dist
Produces dist/com.forkless.mattermore-{version}.tar.gz containing the plugin
manifest and platform binaries.
11.2 Upload
Via Mattermost System Console → Plugins → Upload Plugin, or via API.
11.3 Upgrade
Upload new .tar.gz — Mattermost replaces the plugin binary.
Settings are preserved via KV store. Bot account is re-fetched (not re-created).
11.4 Local Testing
docker compose -f dev/docker-compose.yml up -d
Spins up Mattermost 11.6.1 + PostgreSQL 16 at http://localhost:8065.
Auto-creates admin user (admin/password).
Test rig is disposable: docker compose down -v wipes everything.
12. File Structure
mattermore/
├── plugin.json — Plugin manifest
├── Makefile — Build, dist, clean
├── .golangci.yml — Linter config
├── DESIGN.md — This document
├── PROJECT.md
├── CHANGES.md
├── DECISIONS.md
├── README.md
├── dev/
│ ├── docker-compose.yml — Local Mattermost + PostgreSQL
│ └── init.sh — Admin auto-creation script
├── server/
│ ├── main.go — plugin.ClientMain
│ ├── plugin.go — Plugin struct
│ ├── hooks.go — OnActivate, MessageHasBeenPosted, streaming
│ ├── command.go — /ai and subcommands, persona override
│ ├── configuration.go — Config struct, validation, OnConfigurationChange
│ ├── ollama.go — OllamaClient (streaming + blocking)
│ ├── store.go — ConversationStore (KV-backed)
│ ├── rate.go — RateLimiter (token bucket)
│ ├── engagement.go — EngagementEngine (sleep/active)
│ └── persona.go — PersonaLibrary with 19 personas
└── dist/
└── com.forkless.mattermore-*.tar.gz
13. Future Considerations (v0.2+)
- Multi-modal inputs — Ollama supports images via
/api/chatwithimagesfield. - Tool calling — Ollama supports tool definitions for external APIs.
- Conversation export — Allow users to export AI threads as markdown.
- Streaming into message attachments — Render structured data (tables, JSON, code blocks) with better formatting.