Skip to main content

Developer tools and resources

Build with TritonAI

Personal productivity, automation and building

Use AI for your own work or build a service

For personal productivity, the TritonAI Harness can work across files, browser tasks, and the productivity tools you choose to connect, using reusable skills and memory sources available through the Harness within the permissions you grant. For recurring processes, n8n can connect applications and APIs around a defined trigger. For application building, you can use the Harness and shared APIs to create an integration, then move it into a secure hosting environment supported by IT Services when other people need to rely on it.

The TritonAI service model

  1. 01
    Campus needA real person, a real task, data you are allowed to use.
  2. 02
    Supported build pathStart in the TritonAI Harness, automate with n8n, or use the developer APIs and campus skills.
  3. 03
    Shared AI platformReach approved models through the gateway.
  4. 04
    Owned serviceHosting, support, and review sized to what breaks if it fails.
A campus need moves through a supported build path and the shared AI platform into an owned service.

Start here

Choose the resource you need

01

Browse models

Review available models, how much information each can handle, and current rates.

Open the Model Hub
02

Request access

Get your credentials, agree to the responsible-use terms, and name who owns the project.

Get started

Shared AI platform

Models available through the gateway

The gateway lists these models today, spanning approved enterprise cloud models and UC-hosted open models. The code under each name is the request ID to use through the gateway, and context length is the amount of input a request can carry. Rates and full details stay in the Model Hub.

Models currently listed by the TritonAI gateway with their hosting, type, and context length
ModelHostingTypeContext length
Claude Opus 5
claude-opus-5
Approved enterprise cloudChat and reasoning1M tokens
Claude Sonnet 5
claude-sonnet-5
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.8
claude-opus-4-8
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.7
claude-opus-4-7
Approved enterprise cloudChat and reasoning1M tokens
Claude Sonnet 4.6
claude-sonnet-4-6
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.6
claude-opus-4-6
Approved enterprise cloudChat and reasoning1M tokens
Claude Opus 4.6
claude-opus-4-6-v1
Approved enterprise cloudChat and reasoning200K tokens
Gemini 3.6 Flash
gemini-3.6-flash
Approved enterprise cloudChat and reasoning1.048576M tokens
Gemini 3.5 Flash
gemini-3.5-flash
Approved enterprise cloudChat and reasoning1.048576M tokens
Gemini 3.5 Flash Lite
gemini-3.5-flash-lite
Approved enterprise cloudChat and reasoning1.048576M tokens
Gemma 4 31B
api-gemma-4-31b
UC-hostedChat and reasoning256K tokens
GPT-5.6 Luna
gpt-5.6-luna
Approved enterprise cloudChat and reasoning1.05M tokens
GPT-5.6 Sol
gpt-5.6-sol
Approved enterprise cloudChat and reasoning1.05M tokens
GPT-5.6 Terra
gpt-5.6-terra
Approved enterprise cloudChat and reasoning1.05M tokens
GPT-5.5
gpt-5.5
Approved enterprise cloudChat and reasoning1.05M tokens
GPT-5.4
gpt-5.4
Approved enterprise cloudChat and reasoning1.05M tokens
GPT-OSS 120B
api-gpt-oss-120b
UC-hostedChat and reasoning128K tokens
Kimi K2.6
kimi-k2.6
Approved enterprise cloudChat and reasoningSee Model Hub
Kimi K2.5
moonshotai.kimi-k2.5
Approved enterprise cloudChat and reasoning262K tokens
MiniMax M2
minimax.minimax-m2
Approved enterprise cloudChat and reasoning128K tokens
Mistral Large 3
mistral.mistral-large-3-675b-instruct
Approved enterprise cloudChat and reasoning128K tokens
Amazon Nova 2 Lite
us.amazon.nova-2-lite-v1:0
Approved enterprise cloudChat and reasoning1M tokens
Amazon Nova Premier
us.amazon.nova-premier-v1:0
Approved enterprise cloudChat and reasoning1M tokens
TritonGPT Embeddings
api-tgpt-embeddings
Approved enterprise cloudEmbeddings33K tokens
Tgpt Embeddings
tgpt-embeddings
Approved enterprise cloudEmbeddingsSee Model Hub
DeepSeek V4 Flash
api-deepseek-v4-flash
UC-hostedChat and reasoning1M tokens
GLM 5.2
api-glm-5.2
UC-hostedChat and reasoning320K tokens
LightOn OCR 1B
api-lightonocr-1b
UC-hostedDocument OCR8K tokens
Cohere Transcribe
api-cohere-transcribe
UC-hostedSpeech to textSee Model Hub

List refreshed from the public Model Hub on 2026-08-03. Test registrations and TritonGPT-internal serving entries are excluded.

Primary supported workspace

Use TritonAI Harness for desktop agent work

The Harness is UC San Diego’s main supported desktop workspace for building and running AI agents. Installation, model access, campus skills, and permissions all come set up together, so you are not assembling them yourself. Claude Code and the Codex desktop app are also supported if you already prefer one of those.

Primary supported workspace

TritonAI Harness

A desktop workspace set up for campus, so approved models, skills, and permissions are already wired together.

Supported

Claude Code

A command-line environment, if you would rather work directly with Anthropic models from a terminal.

Supported

Codex desktop app

OpenAI's desktop app, for builders already juggling repositories and parallel agent tasks.

Why we point people at the Harness first

  • One installer for Mac and WindowsBundles the required runtime, packages, skills library, configurations, and the campus-selected Harness version.
  • UC San Diego model accessRoutes requests through the shared gateway to available UC-hosted and approved frontier models.
  • Campus skills built inOfficial skills ship with it, plus a reviewed way to pick up community ones. You do not assemble a toolchain.
  • Managed Microsoft 365 connectionsThe Harness plugin makes structured calls and keeps connection tokens outside the agent context; users choose the permissions they grant.
  • Adjustable supervisionApprove every step, auto-accept, or hand over full access. Set it to whatever the task and your nerves can take.
  • Desktop work in one placeChat, files, images, voice input, and a built-in browser, all on the same approved model route.

Request developer accessBrowse compatible skills

A shared execution path

  1. 01Install onceCampus setup and skills are bundled
  2. 02Connect deliberatelyPick your models, plugins, and permissions
  3. 03Work with contextApproved tools and campus patterns
  4. 04Review the resultYou still answer for the decision

The Harness is an early pilot and changes quickly. What you can access depends on your approved service path.

Workflow automation

Build repeatable workflows with n8n

UC San Diego hosts n8n, a visual workflow-automation platform that connects applications and APIs with little or no code. A workflow can start from a schedule, webhook, email, or file event and then run a defined series of steps. Workflows can also include AI-assisted steps and pause for human review before selected actions.

How n8n fits with the TritonAI Harness

The Harness supports interactive agent work across files, browsers, and connected tools. n8n supports processes that begin from a known trigger and follow a repeatable path. Projects can use the Harness for interactive work and n8n for recurring execution.

Request n8n access Open n8n

The shared API path

Everything goes through one gateway

You build in a supported environment and send model requests to the TritonAI LLM Gateway. One endpoint handles routing to approved cloud or UC-hosted models.

Start with a campus need

Campus builders

  • Department staff
  • Research labs
  • Administrative analysts
  • Faculty teams

Build in a supported workspace

Development environments

  • TritonAI HarnessPrimary supported workspace
  • Claude Code
  • Codex desktop app

Shared managed route

TritonAI
LLM Gateway

Choose an approved route

Model routes

  • Enterprise cloudAWS, Microsoft Azure, and Google Cloud Vertex AI
  • SDSC-hostedLocally hosted at the San Diego Supercomputer Center

Capabilities vary by model

Available capabilities

  • Chat
  • Reasoning
  • Vision
  • Image generation
  • OCR
  • Coding
Getting gateway access does not give your application permission to use new data. Your application is still responsible for approved data, testing, accessibility, support, and human review.

Shared campus compute

TritonAI uses shared campus computing

DataHub is the web front door to the Data Science and Machine Learning Platform (DSMLP), which supplies CPU and GPU capacity, storage, and environments that are ready to go. Coursework, formal independent study, eligible student projects, and some TritonAI workloads all run there.

Explore DSMLP and DataHub

Shared infrastructure at a glance

  1. DataHub and launch toolsWeb and command-line access
  2. DSMLPContainers, compute, storage, and datasets
  3. Selected TritonAI workloadsServices built on shared campus capacity
DataHub and command-line launch tools provide access to DSMLP shared compute, which supports coursework, formal independent study, eligible student projects, and selected TritonAI workloads.

Gateway usage

Gateway usage at a glance

Aggregate TritonAI LLM Gateway activity shows the scale of shared model access across UC San Diego.

  • 309.4Btokens processedTotal input and output tokens recorded by the gateway during the measurement period.
  • 105.1MAPI requestsCompleted gateway request records during the measurement period.
  • 95.3%served by self-hosted modelsShare of recorded tokens served by open-weight models running on UC-controlled infrastructure.
  • 73.2Btokens in JuneHighest monthly token volume recorded in the six-month measurement period.

Monthly token volume

Usage rose to a six-month high in June.

  • Self-hosted
  • Cloud
  1. January44.3B
  2. February49.0B
  3. March44.8B
  4. April50.1B
  5. May48.1B
  6. June73.2B
View monthly data and measurement notes
Gateway token volume by model route, January through June 2026
MonthSelf-hostedCloudTotal tokens
January 202643.8B0.5B44.3B
February 202648.3B0.7B49.0B
March 202643.7B1.0B44.8B
April 202648.5B1.5B50.1B
May 202642.5B5.6B48.1B
June 202667.8B5.4B73.2B
Measurement period
January 1–June 30, 2026
Owner
TritonAI platform team
Data classification
Public aggregate; no user-level records
Last reviewed
2026-07-25
  • Commercial API model families are classified as cloud; open-weight model families running on UC-controlled infrastructure are classified as self-hosted.
  • Approximately 0.15 billion unattributed probe and test tokens are included in the aggregate total but not in audience-level breakdowns.

From request to service

What a prototype needs before it becomes a service

The more people rely on it and the worse the failure, the more of this you have to have in place.

  1. Request

    Name the user, the task, the data you may use, and how you will know it worked.

  2. Prototype

    Keep the data bounded and put a person in the loop on purpose.

  3. Evaluate

    Test quality, accessibility, security, and what it costs to run.

  4. Operate

    Name an owner, write down the controls, support your users, and watch it.

Hosting and support

Pick the hosting lane that matches your reach and risk

Something useful is not yet a production service. As more people depend on it, as it touches more data, or as failure starts to cost something, move it into a more managed lane.

  1. Lane 0Personal workspace
    Best forExplore a bounded task

    Working on your own, in an approved desktop environment or sandbox.

    HostingUser-controlled workspace

    Fine for learning and prototypes. Not something to hand other people.

    URLlocalhost

    AccountabilityIndividual owner

    You protect the data, you check the results, you keep the scope small.

  2. Move to a managed lane as scope grows
  3. Lane 1Department application
    Best forServe a defined team

    One application, a known set of users, and a business owner who wants it.

    HostingDepartment-owned application

    Published through an approved campus application path.

    URLapps.ucsd.edu

    AccountabilityInitial risk and scope review

    Say who maintains it, who answers support, and what data it may touch.

  4. Move to a managed lane as scope grows
  5. Lane 2Managed campus service
    Best forSupport many users

    A shared workflow, usually with integrations, that other teams now depend on.

    HostingTritonAI or ITS-managed path

    A named team operates and supports the service.

    URLtritonai.ucsd.edu

    AccountabilityRecurring risk and scope review

    Teams monitor quality, security, accessibility, and uptime. They also own support.

  6. Move to a managed lane as scope grows
  7. Lane 3Enterprise service
    Best forCampus-wide delivery

    Something the whole university uses, or something that hurts badly when it breaks.

    HostingEnterprise platform

    Built with real architecture, identity, and service management behind it.

    URLucsd.edu

    AccountabilityFormal operating ownership

    Governance, monitoring, continuity, and support come with the service.

Escalate when:
  • Audience or reliance grows
  • Data or integrations expand
  • Failure or support impact rises
Lane is about scope and risk, not headcount. A prototype can move between lanes as what it does and what it needs to run change.

Shared responsibility

Who owns what

Platform responsibilities

We run the gateway, publish the patterns and standards, and review the service path you propose.

Department responsibilities

You own the application logic, the data it uses, accessibility, testing, your users' support, and a named technical owner.

See trust and architecture Read developer FAQs

Before you email us

Bring a narrow, measurable problem

Work out who it is for and what data it may use. Name the review point and a result you can measure. Then choose the model.