Configuring Generative AI
Configuration
A Generative AI provider can be configured in the global config, which will make the Generative AI features available for use. There are currently 5 native providers available to integrate with Nabik. Other providers that support the OpenAI standard API can also be used. See the OpenAI-Compatible section below.
genai is a map of named providers. Each key under genai is a name you choose, and its value is that provider's settings:
- Frigate UI
- YAML
- Navigate to Settings→Enrichments→Generative AI.
- Click Add and enter a Provider name. Any name of letters, numbers, hyphens, and underscores is accepted, but it cannot be changed from the UI after the provider is created.
- Set Provider to the service you are using (e.g.,
ollama) - Set Base URL, API key, and Model as required by that provider
- Set Roles to the roles this provider should handle.
genai:
my_provider: # any name you like
provider: ollama
base_url: http://localhost:11434
model: qwen3-vl:4b
roles:
- descriptions
- embeddings
- chat
The examples on this page all use my_provider, but the name is arbitrary and is only used to reference the provider elsewhere in the config (for example, semantic_search.model).
Each provider handles one or more roles: chat, descriptions, embeddings, and transcribe. A provider handles the first three by default; transcribe must always be listed explicitly, and is not available on Ollama, which has no audio input. Each role may be assigned to exactly one provider. Define a single provider if you want it to do everything, or split the roles across several providers using the roles option.
If the provider you choose requires an API key, you may either directly paste it in your configuration, or store it in an environment variable prefixed with FRIGATE_.
Local Providers
Local providers run on your own hardware and keep all data processing private. These require a GPU or dedicated hardware for best performance.
Running Generative AI models on CPU is not recommended, as high inference times make using Generative AI impractical.
Recommended Local Models
Vision models
You must use a vision-capable model with Nabik. The following models are recommended for local deployment of the descriptions and chat roles:
| Model | Review frame mode | Notes |
|---|---|---|
qwen3-vl | frames | Strong visual and situational understanding, enhanced ability to identify smaller objects and interactions with object. Follows a sequence of frames on its own. |
qwen3.6/qwen3.8 | frames | Strong situational understanding, but missing DeepStack from qwen3-vl leading to worse performance for identifying objects in people's hand and other small details. |
gemma4 | annotated_frames | Strong situational understanding, sometimes resorts to more vague terms like 'interacts' instead of assigning a specific action. Loses track of activity that repeats or reverses, so it benefits from annotated frames. |
Embedding models
The embeddings role needs a different kind of model. Text queries are matched against the stored image embeddings, so the model must be trained to place images and text into the same vector space. A chat or description model will still return vectors when asked, but those vectors are not trained for retrieval and text searches will return poor matches with no error to indicate why.
| Model | Notes |
|---|---|
qwen3-vl-embedding | Multimodal embeddings for Semantic Search. Must be served by llama.cpp started with --embeddings and --mmproj. |
Transcription models
The transcribe role needs a model that accepts audio input. A text-only or vision-only model cannot serve this role. The following are recommended for local deployment of the transcribe role:
| Model | Notes |
|---|---|
qwen3-asr | Dedicated speech recognition model covering 30 languages, and the better choice for transcription quality. It only transcribes, so it cannot be shared with the descriptions or chat roles. |
gemma4 | General multimodal model that accepts audio as well as images, so one served model can cover transcribe alongside the other roles. Transcript quality is below qwen3-asr, particularly on noisy audio. |
Both must be served by llama.cpp started with the matching audio --mmproj. llama.cpp only reports audio support when an audio projector is loaded. Without it Nabik sees the model as text-only and the transcribe role is unavailable in the UI. Nabik transcribes through the server's /v1/audio/transcriptions route, which llama.cpp serves for any audio-capable model.
Each model is available in multiple parameter sizes (3b, 4b, 8b, etc.). Larger sizes are more capable of complex tasks and understanding of situations, but requires more memory and computational resources. It is recommended to try multiple models and experiment to see which performs best.
You should have at least 8 GB of RAM available (or VRAM if running on GPU) to run the 7B models, 16 GB to run the 13B models, and 24 GB to run the 33B models.
Model Types: Instruct vs Thinking
Vision-language models come in instruct variants (fine-tuned to follow instructions and respond concisely), thinking variants (fine-tuned for free-form, speculative reasoning), and hybrid variants that support both modes per request. Most modern vision-language models are hybrid.
Nabik manages reasoning per task automatically:
- Description tasks (object descriptions, review descriptions, review summaries) are synthesis-only and benefit from concise, direct output, so Nabik disables thinking for these calls when the model exposes a per-request toggle.
- Chat lets you toggle thinking on or off from the composer when the configured model supports it.
You can use a pure instruct, hybrid, or thinking-capable model with Nabik. No extra configuration is required to disable thinking for descriptions.
llama.cpp
llama.cpp is a C++ implementation of LLaMA that provides a high-performance inference server.
It is highly recommended to host the llama.cpp server on a machine with a discrete graphics card, or on an Apple silicon Mac for best performance.
Supported Models
You must use a vision capable model with Nabik. The llama.cpp server supports various vision models in GGUF format.
Configuration
All llama.cpp native options can be passed through provider_options, including temperature, top_k, top_p, min_p, repeat_penalty, repeat_last_n, seed, grammar, and more. See the llama.cpp server documentation for a complete list of available parameters.
- Frigate UI
- YAML
- Navigate to Settings