Skip to main content
Chat models use the OpenAI chat completions API.

Z.AI
GLM-5.2
glm-5-2
Parameters: 754B (40B active)Quantization: FP8 (official FP8 release from Z.AI)Context: 384K tokensStrengths: State-of-the-art agentic engineering, long-horizon tool use, sustained reasoning over hundreds of iterationsStructured Outputs: Structured response formatting supportBest for: Agentic engineering tasks, complex coding workflows, repo-level code generation, and long-running tool-use sessionsModel weights: zai-org/GLM-5.2-FP8Configuration repo: tinfoilsh/confidential-glm5-2

Moonshot
Kimi K2.6
kimi-k2-6
Parameters: 1T total (32B activated)Quantization: INT4 weight-only (native quantization-aware release from Moonshot AI)Context: 256K tokensStrengths: Long-horizon coding, image and video understanding, generates code and interfaces from visual inputs, large-scale agent orchestration, strong tool callingStructured Outputs: Structured response formatting supportBest for: Agentic coding, design-to-code workflows, multimodal applications, and long-running tool-based tasks that benefit from strong reasoningModel weights: moonshotai/Kimi-K2.6Configuration repo: tinfoilsh/confidential-kimi-k2-6
Vision + Language: Supports text, image, and video inputs with native reasoning and tool calling for agentic workflows.

Google DeepMind
Gemma 4 31B
gemma4-31b
Parameters: 31BQuantization: None (served in BF16)Context: 256K tokensStrengths: Built-in thinking mode, image understanding, native function calling, multilingual support for 35+ languagesStructured Outputs: Structured response formatting supportBest for: Reasoning tasks, coding, image analysis, and agentic workflows with tool callingModel weights: google/gemma-4-31B-itConfiguration repo: tinfoilsh/confidential-gemma4-31b
Vision + Language: Processes text and image inputs. Features step-by-step reasoning with configurable thinking mode.

OpenAI
GPT-OSS 120B
gpt-oss-120b
Parameters: 117B (5.1B active)Quantization: MXFP4 (native MXFP4 MoE weights, as released by OpenAI)Context: 131K tokensStrengths: Configurable reasoning effort levels, full chain-of-thought access, built-in capabilities including function calling, web browsing, and Python code executionStructured Outputs: Structured response formatting supportBest for: Production use cases requiring configurable reasoning and tool useModel weights: openai/gpt-oss-120bConfiguration repo: tinfoilsh/confidential-gpt-oss-120b

Llama
Llama 3.3 70B
llama3-3-70b
Parameters: 70BQuantization: FP8 (FP8 weights with dynamic FP8 activations, quantized from the BF16 release)Context: 128K tokensStrengths: Multilingual, dialogue-optimized, function callingStructured Outputs: Structured response formatting supportBest for: Conversational AI applications and complex dialogue systemsModel weights: RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamicConfiguration repo: tinfoilsh/confidential-llama3-3-70b