🔀 Model Routing
IntentScope automatically selects the best model for every request based on intent analysis. You never have to hardcode a model name.
Provider Priority
The gateway auto-detects which provider to use based on configured API keys:
| Priority | Provider | Cost | Requires |
|---|---|---|---|
| 1️⃣ | Groq | Free | GROQ_API_KEY |
| 2️⃣ | Mistral AI | Pay-per-token | MISTRAL_API_KEY |
| 3️⃣ | Ollama | Free (local) | Running Ollama instance |
Set GROQ_API_KEY to always use the free Groq tier. If it’s missing, the gateway falls back to Mistral, then Ollama.
Light vs Heavy Tier
Every provider has two model tiers. IntentScope automatically upgrades to the heavy tier for complex or risky tasks.
| Provider | Light Model | Heavy Model |
|---|---|---|
| Groq | llama-3.1-8b-instant | llama-3.3-70b-versatile |
| Mistral | mistral-small-latest | mistral-large-latest |
| Ollama | llama3.2 | llama3.2 |
When does the heavy tier trigger?
needs_upgrade = category in ("code_gen", "analysis") or risk_level == "high"| Intent Category | Risk Level | Tier Used |
|---|---|---|
qa | low | ✅ Light |
summarization | low | ✅ Light |
creative | medium | ✅ Light |
code_gen | any | ⚡ Heavy |
analysis | any | ⚡ Heavy |
| any | high | ⚡ Heavy |
Intent Classification
Before routing, every prompt is classified by intent_analyzer.py using Mistral AI (or Ollama as fallback).
The classifier returns:
{
"category": "code_gen",
"risk_level": "medium",
"estimated_tokens": 800
}Categories:
| Category | Description |
|---|---|
code_gen | Writing or explaining code |
summarization | Condensing or summarizing text |
creative | Writing, storytelling, brainstorming |
data_extraction | Parsing or extracting structured data |
qa | Question answering, factual lookups |
other | Everything else |
Risk levels:
| Level | Triggers |
|---|---|
low | Simple, safe prompts |
medium | Moderate complexity |
high | Potential prompt injection, sensitive data, complex logic |
Routing Decision Flow
Request arrives with model="default"
│
▼
Does user have a saved default_model preference?
│
Yes ─┼─ Is it in KNOWN_VALID_MODELS?
│ │
│ Yes ─→ Use that model as base
│ No ─→ Log warning, fall back to provider default
│
No ─┼─→ Auto-detect provider (Groq → Mistral → Ollama)
│ Use provider's light model as base
│
▼
Is category code_gen/analysis OR risk_level high?
│
Yes ─→ Upgrade to heavy tier
No ─→ Return base modelOverriding the Model
You can pass a specific model in your request to bypass routing:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Authorization: Bearer sk-intentscope-..." \
-d '{
"model": "groq/llama-3.3-70b-versatile",
"messages": [...]
}'Only models in KNOWN_VALID_MODELS are accepted for explicit overrides. Invalid or removed models are ignored and routing falls back to auto-detection.
Adding New Models
To add a new model, update KNOWN_VALID_MODELS and the relevant tier dict in backend-python/app/services/model_router.py:
GROQ_MODELS = {
"light": "groq/llama-3.1-8b-instant",
"heavy": "groq/llama-3.3-70b-versatile",
}
KNOWN_VALID_MODELS: set[str] = {
"groq/llama-3.1-8b-instant",
"groq/llama-3.3-70b-versatile",
"groq/gemma2-9b-it",
# Add new models here
}Then run the DB cleanup SQL to reset any users with stale model preferences:
UPDATE users
SET default_model = 'groq/llama-3.1-8b-instant'
WHERE default_model NOT IN (
'groq/llama-3.1-8b-instant',
'groq/llama-3.3-70b-versatile',
'groq/gemma2-9b-it',
'mistral/mistral-small-latest',
'mistral/mistral-large-latest',
'ollama/llama3.2'
);