The best open LLMs for your use case:

1GLM-5.2Z.ai

Agentic-focused refinement of GLM-5.1 from Z.ai with improved coding, tool use, and reasoning, plus extended 256K context.

Speed:

Intelligence:

Price: (1M Tokens)

$1.40 / 4.40

Cached input: (1M Tokens)

$0.26

Context: (tokens)

262,144

Inputs:

ImageText

Benchmarks:

#1

MCP-Atlas

Agents and Function Calling

76.8
#1

HLE

General Knowledge

40.5
#1

SWE-Bench Pro

Coding Agents

62.1
#1

DeepSWE

Coding Agents

46.2
#1

Terminal-Bench 2.1

Coding Agents

81
#1

FrontierSWE

Coding Agents

74.4
#1

EQBench

Creative Writing

1575
#2

GPQA-Diamond

General Knowledge

91.2
#1

MCP-Atlas

Agents and Function Calling

76.8
#1

HLE

General Knowledge

40.5
#1

SWE-Bench Pro

Coding Agents

62.1
#1

DeepSWE

Coding Agents

46.2
#1

Terminal-Bench 2.1

Coding Agents

81
#1

FrontierSWE

Coding Agents

74.4
#1

EQBench

Creative Writing

1575
#2

GPQA-Diamond

General Knowledge

91.2
2MiniMax-M3MiniMax

Next-generation reasoning model from MiniMax with frontier agentic, coding, and multimodal performance. Strong scores on SWE-Bench, BrowseComp, OmniDocBench, and IMO/USAMO competition reasoning.

Speed:

Intelligence:

Price: (1M Tokens)

$0.30 / 1.20

Cached input: (1M Tokens)

$0.06

Context: (tokens)

524,288

Inputs:

ImageText

Benchmarks:

#2

Apex Agents

Agents and Function Calling

27.7
#1

Claw-Eval

Agents and Function Calling

74.5
#1

GPQA-Diamond

General Knowledge

92.9
#1

Video-MME v2

Multimodal - Vision

85.4
#2

SWE-Bench Verified

Coding Agents

80.5
#2

SWE-Bench Pro

Coding Agents

59
#3

MMMU-Pro

Multimodal - Vision

78.1
#2

Apex Agents

Agents and Function Calling

27.7
#1

Claw-Eval

Agents and Function Calling

74.5
#1

GPQA-Diamond

General Knowledge

92.9
#1

Video-MME v2

Multimodal - Vision

85.4
#2

SWE-Bench Verified

Coding Agents

80.5
#2

SWE-Bench Pro

Coding Agents

59
#3

MMMU-Pro

Multimodal - Vision

78.1

Use case:

Agents and Function Calling

Features:

Low Latency