Skip to main content

Build a service

Build with TritonAI

One key, approved models

Build on the models the campus already approved

The TritonAI LLM Gateway gives every campus builder one key to approved cloud and UC-hosted models. Use it from TritonAI Harness, Claude Code, Codex, Hermes, or any client that connects to the Gateway, from an n8n workflow, or from your own code. When other people start to depend on what you built, move it into hosting and support sized to its users, data, and impact.

Get a Gateway key Plan a first project

What comes with the key

  1. 01
    Campus agreementsCloud models run under UC enterprise contracts. Nothing is trained on your data.
  2. 02
    UC-hosted routesOpen models on campus infrastructure, the no-recharge default for most campus work.
  3. 03
    Approved for P1 to P3The same approval TritonGPT has. P4 data is not approved.
  4. 04
    One place for cost and usageYour approval sets the routes, limits, and billing for the key.
A Gateway key comes with campus enterprise agreements, UC-hosted model routes, approval for Protection Levels 1 through 3, and usage and billing controls set at approval.

Choose your path

How do you want to build?

All three paths use the same Gateway key. Pick the one that matches the work in front of you. You can change paths later as the audience, data, or support needs change.

Teaching a course? DataHub and DSMLP supply course compute and are a separate service from the Gateway.

Cost and eligibility

What a key costs

Recharge is the campus term for internal billing to a department chartstring. Your approval sets which of these apply to your key.

01

Campus work

UC-hosted models carry no recharge for campus administrative work. Monthly caps apply.

02

Cloud models

Billed to a departmental chartstring from the first token at the rate published in the Model Hub. The request names a budget owner and a spend limit.

03

Grant research

UC-hosted and cloud use are charged to the grant or approved project chartstring.

04

Other campuses

Other UC campuses connect through an intercampus recharge agreement arranged with the TritonAI team.

Rates and limits for every route stay in the Model Hub, and the Get Started page walks through the request. The FAQ covers sponsored research, Health Sciences, and other UC campuses in more detail.

Eligibility and setup Funding questions in the FAQ

One API for approved models

Connect through the TritonAI LLM Gateway

Every request from TritonAI Harness, an n8n workflow, or your own code goes through this one endpoint. The Gateway routes each approved key to the models in its approval, and the approval defines access, limits, and billing treatment. The Get Started page covers client setup.

Start with a campus need

Campus builders

  • Department staff
  • Research labs
  • Administrative analysts
  • Faculty teams

Connect a client or application

API clients

  • TritonAI HarnessPrimary supported client
  • Supported alternativesClaude Code and Codex
  • Compatible clientsHermes, OpenCode, and others

Shared API endpoint

TritonAI
LLM Gateway

Choose an approved route

Model routes

  • Enterprise cloudAWS, Microsoft Azure, and Google Cloud Vertex AI
  • UC-hostedOpen models on UC San Diego infrastructure

Capabilities vary by model

Available capabilities

  • Chat
  • Reasoning
  • Vision
  • Image generation
  • OCR
  • Coding
The Gateway key controls model access and limits. The client or application remains responsible for data permissions, testing, accessibility, support, and human review.

Models and routes

Models available through the Gateway

The Gateway lists these models today. UC-hosted open models appear first, followed by approved enterprise cloud models. Start with a UC-hosted route such as GLM 5.3 (api-glm-5.3). Move to a cloud route when a task needs it. The code under each name is the request ID to use through the Gateway, and context length is the amount of input a request can carry. Rates and full details stay in the Model Hub.

Models currently listed by the TritonAI gateway with their hosting, type, and context length
ModelHostingTypeContext length
Gemma 4 31B
api-gemma-4-31b
UC-hostedChat and reasoning256K tokens
GLM 5.3 Flash
api-glm-5.3-flash
UC-hostedChat and reasoning500K tokens
GLM 5.3
api-glm-5.3
UC-hostedChat and reasoning320K tokens
LightOn OCR 1B
api-lightonocr-1b
UC-hostedDocument OCR8K tokens
Cohere Transcribe
api-cohere-transcribe
UC-hostedSpeech to textSee Model Hub
Muse Glimmer 30B
api-muse-glimmer-30b
UC-hostedChat and reasoning262K tokens
OpenAI Privacy Filter
api-openai-privacy-filter
UC-hostedChat and reasoning128K tokens
Claude Opus 5
claude-opus-5
Approved enterprise cloudChat and reasoning1M tokens
Claude Sonnet 5
claude-sonnet-5
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.8
claude-opus-4-8
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.7
claude-opus-4-7
Approved enterprise cloudChat and reasoning1M tokens
Claude Sonnet 4.6
claude-sonnet-4-6
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.6
claude-opus-4-6
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.6
claude-opus-4-6-v1
Approved enterprise cloudChat and reasoning200K tokens
Gemini 3.8 Flash
gemini-3.8-flash
Approved enterprise cloudChat and reasoning1M tokens
Gemini 3.7 Flash
gemini-3.7-flash
Approved enterprise cloudChat and reasoning1M tokens
Gemini 3.6 Flash
gemini-3.6-flash
Approved enterprise cloudChat and reasoning1M tokens
Gemini 3.5 Flash
gemini-3.5-flash
Approved enterprise cloudChat and reasoning1M tokens
Gemini 3.5 Flash Lite
gemini-3.5-flash-lite
Approved enterprise cloudChat and reasoning1M tokens
GPT-5.6 Luna
gpt-5.6-luna
Approved enterprise cloudChat and reasoning1M tokens
GPT-5.6 Sol
gpt-5.6-sol
Approved enterprise cloudChat and reasoning1M tokens
GPT-5.6 Terra
gpt-5.6-terra
Approved enterprise cloudChat and reasoning1M tokens
GPT-5.5
gpt-5.5
Approved enterprise cloudChat and reasoning1M tokens
GPT-5.4
gpt-5.4
Approved enterprise cloudChat and reasoning1M tokens
Kimi K2.6
kimi-k2.6
Approved enterprise cloudChat and reasoningSee Model Hub
Kimi K2.5
moonshotai.kimi-k2.5
Approved enterprise cloudChat and reasoning262K tokens
MiniMax M2
minimax.minimax-m2
Approved enterprise cloudChat and reasoning128K tokens
Mistral Large 3
mistral.mistral-large-3-675b-instruct
Approved enterprise cloudChat and reasoning128K tokens
Amazon Nova 2 Lite
us.amazon.nova-2-lite-v1:0
Approved enterprise cloudChat and reasoning1M tokens
Amazon Nova Premier
us.amazon.nova-premier-v1:0
Approved enterprise cloudChat and reasoning1M tokens
TritonGPT Embeddings
api-tgpt-embeddings
Approved enterprise cloudEmbeddings4K tokens

List refreshed from the public Model Hub on 2026-09-21. Test registrations and TritonGPT-internal serving entries are excluded.

Choose a client

TritonAI Harness and other clients

TritonAI Harness is UC San Diego's primary supported client. It is in pilot, runs on Mac (Apple Silicon) and Windows, and anyone with a Gateway key can request it. Claude Code and Codex are supported alternatives. Other compatible clients can connect with the same endpoint and key, though their features and setup differ.

Primary supported client

TritonAI Harness

A desktop workspace with the Gateway connection, campus skills, and Microsoft 365, Google Workspace, and GitHub connections set up for UC San Diego use.

Supported alternatives

Claude Code and Codex

Keep a terminal or desktop workflow you already use and point it at the model routes approved for your key.

Compatible clients

Hermes, OpenCode, and others

Connect with the Gateway endpoint and key from your approval. Setup and support are yours.

Explore TritonAI Harness Download and set up

Workflow automation

Build repeatable workflows with n8n

UC San Diego hosts n8n, a visual workflow platform that connects applications and APIs with little or no code. A workflow starts from a schedule, webhook, email, or file event and runs a defined series of steps. Model requests inside a workflow go through the Gateway with your key, and a workflow can pause for a person before selected actions.

n8n fits best once you know the process and how it should handle exceptions. For work that changes shape every time, start in TritonAI Harness.

Request n8n access Open n8n

Built on TritonAI

What campus teams have built

These three run in production today. Each one started as a bounded campus problem with a named owner, and each keeps a person checking the results. Start with the one that looks most like your problem.

Weekly planner grid of color-coded course sections compared across schedule options before enrollment
Production

Class Planner App

Undergraduates build conflict-free schedule options from live section data, compare them side by side, and hand the finished plan to TSS for booking.

Measure: Schedule drafts created, alternatives compared, preference constraints honored, and handoffs to TSS completed.

Visitor checking in at a kiosk before following a privacy-safe queue path to a staffed service counter
Production

Passport Visitor Management

Visitors check in for Passport Services at CSC or the UCSD Bookstore while staff manage each location from a shared queue dashboard.

Measure: Check-in completion, wait-time visibility, queue accuracy, staff workflow efficiency, and service reliability.

Staff hands arranging users, data boundaries, risks, and outcomes on a collaborative use-case canvas
Production

AI Use-Case Meeting

Biweekly sessions where campus staff bring an AI idea and leave with a scoped use case and a recommendation on whether to proceed.

Measure: Time to decision, completeness of intake, appropriate routing, and participant usefulness.

View all use cases

Gateway usage

Gateway usage at a glance

Aggregate TritonAI LLM Gateway activity shows the scale of shared model access across UC San Diego.

  • 455.3Btokens processedTotal input and output tokens recorded by the gateway during the measurement period.
  • 143.9MAPI requestsGateway request records during the measurement period, including successful and failed requests.
  • 92.6%self-hosted and internal routesShare of recorded tokens not classified as commercial cloud, including internal TritonAI routes.
  • 78.9Btokens in AugustTotal token volume recorded in the latest complete month.

Monthly token volume

August was the highest-volume month to date.

  • Self-hosted and internal
  • Cloud
  1. January44.3B
  2. February49.0B
  3. March44.8B
  4. April50.1B
  5. May48.1B
  6. June73.2B
  7. July67.1B
  8. August78.9B
View monthly data and measurement notes
Gateway token volume by model route, January 1–August 31, 2026
MonthSelf-hosted and internalCloudTotal tokens
January 202643.8B0.5B44.3B
February 202648.3B0.7B49.0B
March 202643.7B1.0B44.8B
April 202648.7B1.4B50.1B
May 202642.8B5.3B48.1B
June 202668.3B4.9B73.2B
July 202661.5B5.6B67.1B
August 202664.5B14.3B78.9B

Monthly drivers

Month-to-month changes reflect the mix of campus applications, automated services, model routing, and instructional activity using the TritonAI Gateway.

  • January: 44.3B tokens

    Usage was driven primarily by embedding and automated background-processing workloads, establishing the baseline for the reporting period.

  • February: 49.0B tokens

    Usage increased as application, widget, and shared API activity expanded across the platform.

  • March: 44.8B tokens

    Usage moderated as widget and automated processing activity declined, partially offset by growth through newer model routes.

  • April: 50.1B tokens

    Usage rebounded with increased embedding, background-processing, and application activity.

  • May: 48.1B tokens

    Overall usage remained relatively steady while workloads shifted toward newer self-hosted and cloud model routes.

  • June: 73.2B tokens

    June reached 73.2 billion tokens, driven by broad growth across production applications, embeddings, and automated services.

  • July: 67.1B tokens

    Usage remained elevated but declined from June as application requests, embeddings, and background processing slowed. Growth through newer model routes offset part of the reduction. Lower summer instructional activity accounted for approximately 0.8B tokens, or 12.7% of the decline, so most of the reduction came from application and platform workloads.

  • August: 78.9B tokens

    August reached the highest monthly volume of the year to date at 78.9 billion tokens, up 17.6% from July. Cloud usage increased to 14.3 billion tokens, while self-hosted and internal routes accounted for 64.5 billion tokens.

Measurement period
January 1–August 31, 2026
Owner
TritonAI platform team
Data classification
Public aggregate; no user-level records
Last reviewed
2026-09-03
  • January through July retains the reviewed V3 model-route classification. August uses the reviewed classification with the September 2 hosting-audit correction, moving 8,340,260 tokens from self-hosted to cloud. Self-hosted and internal includes all remaining recorded tokens, including internal TritonAI routes.
  • The monthly series reconciles to 455,295,685,475 tokens and 143,909,714 request records for January through August 2026.
  • Token totals use the gateway's recorded input plus output tokens. August cache-read and cache-creation fields are not added separately. The earlier January-July export did not expose those fields, so an upstream token-definition change cannot be completely ruled out.
  • August includes all 31 days and 18,971,745 request records. Daily token and request totals reconcile exactly between the supplied normalized export and raw gateway response.

From prototype to service

Add hosting and support as more people rely on it

Something useful is not yet a service. As more people depend on it, as it touches more data, or as failure starts to cost something, move it up a rung. The strategy page describes the program lifecycle behind this ladder.

  1. Rung 1Prototype
    Best forTesting an idea

    Working on your own with sample data you are approved to use.

    HostingYour own workspace

    TritonAI Harness or a local sandbox. Fine for learning. Not something to hand other people.

    AccountabilityYou

    You protect the data, check the results, and keep the scope small.

  2. Rung 2Team workflow
    Best forA recurring job for a known group

    One application or workflow, a defined set of users, and a business owner who wants it.

    HostingDepartment-owned application or n8n workflow

    Published through an approved campus application path with campus sign-in.

    AccountabilityInitial risk and scope review

    Say who maintains it, who answers support, and what data it may touch.

  3. Rung 3Campus service
    Best forServing people outside your team

    A shared workflow, usually with integrations, that other units now depend on.

    HostingTritonAI or ITS-managed path

    A named team operates and supports the service.

    AccountabilityRecurring review

    A named team monitors quality, security, accessibility, and uptime, and owns support.

  4. Rung 4Enterprise service
    Best forCampus-wide delivery

    Something the whole university uses, or something that hurts badly when it breaks.

    HostingEnterprise platform

    Architecture, identity, and service management behind it.

    AccountabilityFormal operating ownership

    Governance, monitoring, continuity, and support come with the service.

Move up when:
  • Audience or reliance grows
  • Data or integrations expand
  • Failure or support impact rises
The platform team runs the Gateway, publishes the patterns, and reviews the service path you propose. Your department owns the application, its data, accessibility, testing, user support, and a named technical owner. See who owns what.

Get a Gateway key

Request access and run a first test

The Get Started page covers eligibility, funding, key protection, client choice, and installation. Most requests need only the form and a short description of the task.