The best open LLMs for your use case:
DeepSeek's frontier 1.6T-parameter Mixture-of-Experts model (49B active per token) with hybrid attention built for long-context, low-cost reasoning. Runs in FP4 with a 512K-token context window.
Speed:
Intelligence:
Price: (1M Tokens)
$1.74 / 3.48Cached input: (1M Tokens)
$0.20Context: (tokens)
512,000Inputs:
Benchmarks:
GPQA-Diamond
General Knowledge
MMLU-Pro
General Knowledge
HLE
General Knowledge
SimpleQA
General Knowledge
LiveCodeBench
Coding Agents
SWE-Bench Verified
Coding Agents
Terminal-Bench 2.0
Coding Agents
GDPval-AA
Agents and Function Calling
GPQA-Diamond
General Knowledge
MMLU-Pro
General Knowledge
HLE
General Knowledge
SimpleQA
General Knowledge
LiveCodeBench
Coding Agents
SWE-Bench Verified
Coding Agents
Terminal-Bench 2.0
Coding Agents
GDPval-AA
Agents and Function Calling
1T-parameter MoE flagship from Moonshot with long-horizon coding, agent swarms scaling to 300 sub-agents, and state-of-the-art reasoning.
Speed:
Intelligence:
Price: (1M Tokens)
$1.20 / 4.50Cached input: (1M Tokens)
$0.20Context: (tokens)
262,144Inputs:
Benchmarks:
GPQA-Diamond
General Knowledge
HLE
General Knowledge
SciCode
Coding Agents
MCP-Mark
Agents and Function Calling
MMMU-Pro
Multimodal - Vision
Apex Agents
Agents and Function Calling
FrontierCode
Coding Agents
LiveCodeBench
Coding Agents
GPQA-Diamond
General Knowledge
HLE
General Knowledge
SciCode
Coding Agents
MCP-Mark
Agents and Function Calling
MMMU-Pro
Multimodal - Vision
Apex Agents
Agents and Function Calling
FrontierCode
Coding Agents
LiveCodeBench
Coding Agents
Use case:
General Knowledge
Features:
Long Context Handling