AI Models
AI Models registers the LLM endpoints the organization can use and the model configurations approved on each one.
Navigation: Admin hub > AI Models (or go to /admin/ai-models)
The page heading is Endpoints and Models.
Identical tool-call loop guard
Identical tool-call loop guard is a cluster-wide limit: how many identical tool calls (the same tool and the same arguments) an agent may make before its loop is stopped. It applies to every model. Change the number and click Save. A toast confirms Loop guard saved.
Add an endpoint
- Click Add endpoint.
- In the Add endpoint dialog:
- Host — pick a host already provisioned for this deployment. You do not type the URL or a credential. The list is empty until a host has been provisioned.
- Name — a label for this endpoint (for example,
GenAI.mil gateway). - Enabled — turn the endpoint off to keep it registered without using it.
- Click Add endpoint. A toast confirms Endpoint added. Editing an existing endpoint confirms Endpoint saved.
The same host can back more than one endpoint. When it does, the host label is disambiguated.
An endpoint card shows the name, the host, and a Disabled badge when the endpoint is off. Use the pencil to open Edit endpoint, or the trash icon to delete it. Deleting asks Delete {name}? and, after you confirm, toasts Deleted {name}. Deleting an endpoint also removes every model configured on it.
Add a model
On an endpoint card, click Add model.
In the Add model dialog:
- Configuration name — the label people see (for example,
Gemini 3.6 Flash (Path B)). The same model can be added more than once with different delivery settings. - Model ID — the identifier passed to the LLM client (for example,
gemini/gemini-3.6-flash). The provider is inferred from this identifier. - Path — URL path appended to the endpoint's host (for example,
/v1). The dialog shows the effective URL. - Disable system prompt — fold the grounding into the first user message instead of a system role.
- Disable native tool-calling interface — drive tools with a prompt-based loop instead of the provider's native tool calls.
- Show in the user model picker — let people choose this model in chat. If no models are selectable, the picker stays hidden.
- Max iterations — the tool-loop cap for this model. Leave it blank to use the cluster default.
- Extended thinking — turn on thinking for this model, then set Thinking budget (tokens).
- Reasoning effort — the default effort sent for this model. No override leaves it unset. The choices are minimal, low, medium, and high.
- Reasoning effort choices offered to users — which of those efforts people can pick. Leave them all unchecked to hide the effort picker for this model.
- Enabled — turn the model off to keep the configuration without offering it.
Click Add model. A toast confirms Model added. Editing an existing model confirms Model saved. Cancel closes the dialog without saving.
A model row shows the configuration name, the model ID, and badges for the settings that are on (No system prompt, No native tools, User-selectable, Max iterations). Use the pencil to open Edit model, or the trash icon to delete that configuration only. Delete uses the same Delete {name}? confirmation and Deleted {name}. toast.
Choose the default model
The default model is the one each new chat starts on. It shows a Default badge. On any other enabled model, click Set as default. A toast confirms Default model updated. A disabled model cannot be the default.