Skip to content
TechCrunch·

🚨GLM-5.2 Closes Gap With GPT-5.6 Sol and Claude Opus

How close are Chinese models to OpenAI's?

TL;DR

SaferAI finds GLM-5.2 just months behind industry leaders on cyber and bio tasks, raising questions about risk management for open-weight models.

GLM-5.2, a Chinese open-weight model, has nearly caught up to OpenAI's GPT-5.6 Sol and Anthropic's Claude Opus in cyber and bio capabilities. This shift means the debate is now focused on managing risks rather than just competing. Developers must consider how these models might be misused once they're freely available. GLM-5.2 refused none of the offensive tasks it was given, while Claude Opus 4.7 consistently refused such requests, making it harder to complete CyberGym tasks.

GLM-5.2 Closes Gap With GPT-5.6 Sol and Claude Opus — TechCrunch

Key Points

1

SaferAI’s evaluation shows GLM-5.2 just months behind GPT-5.6 Sol and Claude Opus 4.7 on cyber and bio tasks

2

GLM-5.2 refused no offensive cyber or dual-use biology tasks, unlike Anthropic's consistent refusal policy

3

Chinese models like GLM-5.2 are designed to run with any safeguards, making misuse harder to prevent

4

Frontier developers rely on classifiers, refusal training, and API-level controls for risk mitigation

5

Data filtering is suggested but impractical for cybersecurity; selective restrictions are preferred

Why It Matters

If you're developing AI models or managing security risks, GLM-5.2's performance highlights the need to reassess how open-weight models can be misused. The gap closing means potential attackers have access to highly capable systems with fewer barriers.

AIGLM-5.2GPT-5.6 SolClaude OpusCyber Security

Frequently Asked Questions

Why does this matter?

If you're developing AI models or managing security risks, GLM-5.2's performance highlights the need to reassess how open-weight models can be misused. The gap closing means potential attackers have access to highly capable systems with fewer barriers.

What happened?

SaferAI finds GLM-5.2 just months behind industry leaders on cyber and bio tasks, raising questions about risk management for open-weight models.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.