Browse models
Review available models, how much information each can handle, and current rates.
Open the Model HubPersonal productivity, automation and building
For personal productivity, the TritonAI Harness can work across files, browser tasks, and the productivity tools you choose to connect, using reusable skills and memory sources available through the Harness within the permissions you grant. For recurring processes, n8n can connect applications and APIs around a defined trigger. For application building, you can use the Harness and shared APIs to create an integration, then move it into a secure hosting environment supported by IT Services when other people need to rely on it.
The TritonAI service model
Start here
Review available models, how much information each can handle, and current rates.
Open the Model HubGet your credentials, agree to the responsible-use terms, and name who owns the project.
Get startedCheck whether another campus team has already written the skill you need.
Browse the Skills LibraryShared AI platform
The gateway lists these models today, spanning approved enterprise cloud models and UC-hosted open models. The code under each name is the request ID to use through the gateway, and context length is the amount of input a request can carry. Rates and full details stay in the Model Hub.
| Model | Hosting | Type | Context length |
|---|---|---|---|
Claude Opus 5claude-opus-5 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Claude Sonnet 5claude-sonnet-5 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Claude Opus 4.8claude-opus-4-8 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Claude Opus 4.7claude-opus-4-7 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Claude Sonnet 4.6claude-sonnet-4-6 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Claude Opus 4.6claude-opus-4-6 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Claude Opus 4.6claude-opus-4-6-v1 | Approved enterprise cloud | Chat and reasoning | 200K tokens |
Gemini 3.6 Flashgemini-3.6-flash | Approved enterprise cloud | Chat and reasoning | 1.048576M tokens |
Gemini 3.5 Flashgemini-3.5-flash | Approved enterprise cloud | Chat and reasoning | 1.048576M tokens |
Gemini 3.5 Flash Litegemini-3.5-flash-lite | Approved enterprise cloud | Chat and reasoning | 1.048576M tokens |
Gemma 4 31Bapi-gemma-4-31b | UC-hosted | Chat and reasoning | 256K tokens |
GPT-5.6 Lunagpt-5.6-luna | Approved enterprise cloud | Chat and reasoning | 1.05M tokens |
GPT-5.6 Solgpt-5.6-sol | Approved enterprise cloud | Chat and reasoning | 1.05M tokens |
GPT-5.6 Terragpt-5.6-terra | Approved enterprise cloud | Chat and reasoning | 1.05M tokens |
GPT-5.5gpt-5.5 | Approved enterprise cloud | Chat and reasoning | 1.05M tokens |
GPT-5.4gpt-5.4 | Approved enterprise cloud | Chat and reasoning | 1.05M tokens |
GPT-OSS 120Bapi-gpt-oss-120b | UC-hosted | Chat and reasoning | 128K tokens |
Kimi K2.6kimi-k2.6 | Approved enterprise cloud | Chat and reasoning | See Model Hub |
Kimi K2.5moonshotai.kimi-k2.5 | Approved enterprise cloud | Chat and reasoning | 262K tokens |
MiniMax M2minimax.minimax-m2 | Approved enterprise cloud | Chat and reasoning | 128K tokens |
Mistral Large 3mistral.mistral-large-3-675b-instruct | Approved enterprise cloud | Chat and reasoning | 128K tokens |
Amazon Nova 2 Liteus.amazon.nova-2-lite-v1:0 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
Amazon Nova Premierus.amazon.nova-premier-v1:0 | Approved enterprise cloud | Chat and reasoning | 1M tokens |
TritonGPT Embeddingsapi-tgpt-embeddings | Approved enterprise cloud | Embeddings | 33K tokens |
Tgpt Embeddingstgpt-embeddings | Approved enterprise cloud | Embeddings | See Model Hub |
DeepSeek V4 Flashapi-deepseek-v4-flash | UC-hosted | Chat and reasoning | 1M tokens |
GLM 5.2api-glm-5.2 | UC-hosted | Chat and reasoning | 320K tokens |
LightOn OCR 1Bapi-lightonocr-1b | UC-hosted | Document OCR | 8K tokens |
Cohere Transcribeapi-cohere-transcribe | UC-hosted | Speech to text | See Model Hub |
List refreshed from the public Model Hub on 2026-08-03. Test registrations and TritonGPT-internal serving entries are excluded.
Primary supported workspace
The Harness is UC San Diego’s main supported desktop workspace for building and running AI agents. Installation, model access, campus skills, and permissions all come set up together, so you are not assembling them yourself. Claude Code and the Codex desktop app are also supported if you already prefer one of those.
Primary supported workspace
A desktop workspace set up for campus, so approved models, skills, and permissions are already wired together.
Supported
A command-line environment, if you would rather work directly with Anthropic models from a terminal.
Supported
OpenAI's desktop app, for builders already juggling repositories and parallel agent tasks.
A shared execution path
The Harness is an early pilot and changes quickly. What you can access depends on your approved service path.
Workflow automation
UC San Diego hosts n8n, a visual workflow-automation platform that connects applications and APIs with little or no code. A workflow can start from a schedule, webhook, email, or file event and then run a defined series of steps. Workflows can also include AI-assisted steps and pause for human review before selected actions.
The Harness supports interactive agent work across files, browsers, and connected tools. n8n supports processes that begin from a known trigger and follow a repeatable path. Projects can use the Harness for interactive work and n8n for recurring execution.
The shared API path
You build in a supported environment and send model requests to the TritonAI LLM Gateway. One endpoint handles routing to approved cloud or UC-hosted models.
Start with a campus need
Build in a supported workspace
Shared managed route
Choose an approved route
Capabilities vary by model
Gateway usage
Aggregate TritonAI LLM Gateway activity shows the scale of shared model access across UC San Diego.
Usage rose to a six-month high in June.
| Month | Self-hosted | Cloud | Total tokens |
|---|---|---|---|
| January 2026 | 43.8B | 0.5B | 44.3B |
| February 2026 | 48.3B | 0.7B | 49.0B |
| March 2026 | 43.7B | 1.0B | 44.8B |
| April 2026 | 48.5B | 1.5B | 50.1B |
| May 2026 | 42.5B | 5.6B | 48.1B |
| June 2026 | 67.8B | 5.4B | 73.2B |
From request to service
The more people rely on it and the worse the failure, the more of this you have to have in place.
Name the user, the task, the data you may use, and how you will know it worked.
Keep the data bounded and put a person in the loop on purpose.
Test quality, accessibility, security, and what it costs to run.
Name an owner, write down the controls, support your users, and watch it.
Hosting and support
Something useful is not yet a production service. As more people depend on it, as it touches more data, or as failure starts to cost something, move it into a more managed lane.
Working on your own, in an approved desktop environment or sandbox.
Fine for learning and prototypes. Not something to hand other people.
URLlocalhost
You protect the data, you check the results, you keep the scope small.
One application, a known set of users, and a business owner who wants it.
Published through an approved campus application path.
URLapps.ucsd.edu
Say who maintains it, who answers support, and what data it may touch.
A shared workflow, usually with integrations, that other teams now depend on.
A named team operates and supports the service.
URLtritonai.ucsd.edu
Teams monitor quality, security, accessibility, and uptime. They also own support.
Something the whole university uses, or something that hurts badly when it breaks.
Built with real architecture, identity, and service management behind it.
URLucsd.edu
Governance, monitoring, continuity, and support come with the service.
Before you email us
Work out who it is for and what data it may use. Name the review point and a result you can measure. Then choose the model.