Model Capabilities
Skulk supports a wide range of models, but not every model behaves the same way.
Some models need:
- custom prompt rendering
- non-generic reasoning delimiters
- specialized tool-call formats
- native multimodal execution
- model-specific API controls
The model capability system exists so Skulk can support those differences without turning the runtime into a pile of hidden one-off checks.
The Two Layers
Skulk now treats capability handling as two related layers:
1. Declarative model card
The model card stores broad static metadata plus optional advanced capability sections.
This is the durable, syncable, editable source of truth.
2. Resolved runtime profile
At runtime, Skulk resolves the card plus tokenizer/model-family facts into a normalized capability profile.
That resolved profile answers questions like:
- should this request use a custom prompt renderer?
- what reasoning format should be expected?
- which output parser should run?
- what defaults should be used when thinking is toggled on or off?
Why Not Only One Layer?
If Skulk used only model cards directly at runtime:
- execution code would be full of
Nonechecks and partial fallbacks - every hot path would need to re-interpret optional metadata
- backward compatibility would be harder to preserve cleanly
If Skulk used only hard-coded runtime profiles:
- custom cards would not be expressive enough
- API and dashboard metadata would drift away from runtime behavior
- model support would become scattered again
The combined approach gives us:
- one declarative source of truth
- one normalized execution contract
The capability spine
The capability system is the spine that UI and API behavior depend on:
- cards can declare advanced capability sections
- older cards without those sections still work
- runtime behavior for key decisions is capability-driven
- model metadata exposed by the API surfaces refined behavior to clients
The decisions it drives today are:
- reasoning/thinking defaults
- prompt renderer selection
- output parser selection
- speech model discovery and TTS/STT dashboard affordances
Thinking contract
The existing public controls are preserved:
enable_thinkingreasoning_effort
But their behavior is now explicitly model-aware through resolved_capabilities.
Toggleable reasoning models
If resolved_capabilities.supports_thinking_toggle is true:
enable_thinking=trueenables thinking using the model profile's default effort unless an explicit non-disabled effort is providedenable_thinking=falsedisables thinking using the profile's disabled effort- omitting both
enable_thinkingandreasoning_effortdisables thinking using the profile's disabled effort reasoning_effort="none"also disables thinking
Non-toggleable reasoning models
If a model supports reasoning but does not support thinking toggle:
- clients should not offer a toggle
- explicit toggle overrides are normalized away
- explicit non-disabled
reasoning_effortvalues are still preserved - requests otherwise fall back to the model's supported default behavior
This keeps the public API stable without pretending every reasoning-capable model can switch on and off cleanly.
Speech contract
Speech models are represented as first-class model-card tasks instead of model name conventions:
TextToSpeechSpeechToTextSpeechTranslation
Cards can add an [audio] section to describe speech-specific behavior:
[audio]
kind = "tts" # or "stt"
default_response_format = "mp3"
response_formats = ["mp3", "wav"]
supports_streaming = true
supports_realtime = false
supports_voice_listing = true
supports_reference_audio = false
supports_translation = false
sample_rates = [16000, 24000]
The resolver exposes that as resolved_capabilities fields:
supports_speech_synthesissupports_transcriptionsupports_speech_translationsupports_audio_outputsupports_realtime_audiodefault_audio_response_formataudio_response_formats
These fields are now runtime-facing metadata. They let placement route speech
cards to the mlx_audio runner, let /v1/models identify mounted TTS/STT
models, and let the dashboard expose voice controls without guessing from model
names. Mounted supports_speech_synthesis models serve /v1/audio/speech.
Cards with a fixed speaker inventory may declare default_voice; Skulk applies
it only when the caller omits voice, and schema validation requires it to be
one of the card's voices.
When the card declares audio.supports_streaming = true, clients can pass
stream=true for stable chunked HTTP MP3 output; bundled cards keep that flag
off until a real MLX model has passed streaming validation. Mounted
supports_transcription models serve
non-streaming /v1/audio/transcriptions.
Cards that additionally declare both streaming and realtime support can expose
the stable stt.realtime@1.0.0 bidirectional provider when the API can reach a
ready single-host runner. The provider accepts mono PCM16, requires a true
upstream incremental session, and does not infer realtime support from a batch
transcription API.
Fallback Behavior
If a model card does not define advanced sections, Skulk should still work.
The runtime resolves that model to a conservative generic profile:
- generic prompt rendering
- generic parser behavior
- no assumptions about special reasoning controls
- no assumptions about special modalities or tool grammars
This is critical for compatibility with existing built-in and custom cards.
Precedence Rules
The resolved runtime profile follows a simple precedence model:
- explicit advanced fields from the model card win
- model-family defaults fill in known behavior for important families
- generic fallback preserves compatibility for everything else
The runtime intentionally keeps those heuristics conservative. The goal is not to guess every possible advanced feature, but to preserve current behavior while letting extended cards make support more precise.
That same approach now applies to builtin platform tools such as web_search:
cards can declare the tool contract, while exposure and execution stay
deployment-aware and family-specific.
What This Enables
Once the capability spine exists, Skulk can evolve cleanly toward:
- model-aware thinking controls
- reasoning budget support
- speech serving controls for TTS, transcription, and translation
- richer tool grammars
- safer dashboard controls based on real support instead of guesswork