Skip to main content
MoonshotMoonshot
Kimi K3
kimi-k3
Parameters: 2.8T total (104B activated)Quantization: MXFP4 weights / MXFP8 activations (native quantization-aware release from Moonshot AI)Context: 256K tokensStrengths: Native image understanding, visual reasoning, document analysis, screenshot-to-code generation, and vision-guided agentic workflowsBest for: Visual knowledge work, document understanding, design-to-code workflows, and multimodal agentsModel weights: moonshotai/Kimi-K3Configuration repo: tinfoilsh/confidential-kimi-k3
Multimodal: Supports text and image inputs with native reasoning and tool calling. See Image Processing Guide for usage examples.
Z.AI
GLM-5.3 Flash
glm-5-3-flash
Parameters: 320B (18B active)Quantization: NVFP4 transformer linear layers (Red Hat AI release; vision tower, embeddings, and output head kept in original precision)Context: 1M tokensStrengths: Image understanding and document analysis over very long context, with reasoning and tool callingBest for: Multimodal document workflows, image analysis alongside large text inputs, and vision-guided agentic workflowsModel weights: RedHatAI/GLM-5.3-Flash-NVFP4Configuration repo: tinfoilsh/confidential-glm5-3-flash
Multimodal: Supports text and image inputs with always-on reasoning and tool calling. See Image Processing Guide for usage examples.
DeepSeek
DeepSeek V4.1 Flash
deepseek-v4-1-flash
Experimental: this model is not yet fully productionized. Capacity and performance may change while we finish rolling it out.
Parameters: 552B (MoE, 6 of 384 routed experts plus 1 shared active per token; Engram memory)Quantization: FP8 attention / FP4 experts (native quantized release from DeepSeek)Context: 1M tokensStrengths: Image understanding over very long context, optional reasoning with configurable effort, built-in speculative decoding for fast generation, tool callingBest for: Image-grounded document analysis alongside very large text inputs, and vision-guided agentic workloadsModel weights: deepseek-ai/DeepSeek-V4.1-FlashConfiguration repo: tinfoilsh/confidential-deepseek-v4-1-flash
Multimodal: Supports text and image inputs with optional reasoning and tool calling. See Image Processing Guide for usage examples.
Google DeepMind
Gemma 4 31B
gemma4-31b
Parameters: 31BQuantization: None (served in BF16)Context: 256K tokensStrengths: Image understanding, object detection, document parsing, OCR, chart comprehension, and pointingBest for: Image analysis, document understanding, OCR tasks, and visual reasoning with built-in thinking modeModel weights: google/gemma-4-31B-itConfiguration repo: tinfoilsh/confidential-gemma4-31b
Multimodal: Supports variable aspect ratios and configurable image token budgets for balancing speed and detail. See Image Processing Guide for usage examples.