Model Routing

🔀 Model Routing

IntentScope automatically selects the best model for every request based on intent analysis. You never have to hardcode a model name.

Provider Priority

The gateway auto-detects which provider to use based on configured API keys:

PriorityProviderCostRequires
1️⃣GroqFreeGROQ_API_KEY
2️⃣Mistral AIPay-per-tokenMISTRAL_API_KEY
3️⃣OllamaFree (local)Running Ollama instance

Set GROQ_API_KEY to always use the free Groq tier. If it’s missing, the gateway falls back to Mistral, then Ollama.

Light vs Heavy Tier

Every provider has two model tiers. IntentScope automatically upgrades to the heavy tier for complex or risky tasks.

ProviderLight ModelHeavy Model
Groqllama-3.1-8b-instantllama-3.3-70b-versatile
Mistralmistral-small-latestmistral-large-latest
Ollamallama3.2llama3.2

When does the heavy tier trigger?

needs_upgrade = category in ("code_gen", "analysis") or risk_level == "high"
Intent CategoryRisk LevelTier Used
qalow✅ Light
summarizationlow✅ Light
creativemedium✅ Light
code_genany⚡ Heavy
analysisany⚡ Heavy
anyhigh⚡ Heavy

Intent Classification

Before routing, every prompt is classified by intent_analyzer.py using Mistral AI (or Ollama as fallback).

The classifier returns:

{
  "category": "code_gen",
  "risk_level": "medium",
  "estimated_tokens": 800
}

Categories:

CategoryDescription
code_genWriting or explaining code
summarizationCondensing or summarizing text
creativeWriting, storytelling, brainstorming
data_extractionParsing or extracting structured data
qaQuestion answering, factual lookups
otherEverything else

Risk levels:

LevelTriggers
lowSimple, safe prompts
mediumModerate complexity
highPotential prompt injection, sensitive data, complex logic

Routing Decision Flow

Request arrives with model="default"
        │
        ▼
Does user have a saved default_model preference?
        │
   Yes ─┼─ Is it in KNOWN_VALID_MODELS?
        │       │
        │    Yes ─→ Use that model as base
        │    No  ─→ Log warning, fall back to provider default
        │
   No  ─┼─→ Auto-detect provider (Groq → Mistral → Ollama)
        │    Use provider's light model as base
        │
        ▼
Is category code_gen/analysis OR risk_level high?
        │
   Yes ─→ Upgrade to heavy tier
   No  ─→ Return base model

Overriding the Model

You can pass a specific model in your request to bypass routing:

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer sk-intentscope-..." \
  -d '{
    "model": "groq/llama-3.3-70b-versatile",
    "messages": [...]
  }'
⚠️

Only models in KNOWN_VALID_MODELS are accepted for explicit overrides. Invalid or removed models are ignored and routing falls back to auto-detection.

Adding New Models

To add a new model, update KNOWN_VALID_MODELS and the relevant tier dict in backend-python/app/services/model_router.py:

GROQ_MODELS = {
    "light": "groq/llama-3.1-8b-instant",
    "heavy": "groq/llama-3.3-70b-versatile",
}
 
KNOWN_VALID_MODELS: set[str] = {
    "groq/llama-3.1-8b-instant",
    "groq/llama-3.3-70b-versatile",
    "groq/gemma2-9b-it",
    # Add new models here
}

Then run the DB cleanup SQL to reset any users with stale model preferences:

UPDATE users
SET default_model = 'groq/llama-3.1-8b-instant'
WHERE default_model NOT IN (
    'groq/llama-3.1-8b-instant',
    'groq/llama-3.3-70b-versatile',
    'groq/gemma2-9b-it',
    'mistral/mistral-small-latest',
    'mistral/mistral-large-latest',
    'ollama/llama3.2'
);