Briefing

GLM Shows Schizophrenic Behavior in Plan‑Mode and Tool‑Calling Benchmarks

ai-dev
by /u/No_Run8812 ·

Benchmark GLM's plan mode and tool calling; it achieves 69% plan_mode success, 71% plan_mode_stress, 90% tool_calling, 67% file_generation, and 75% combined, lower than Qwen3-coder-next.

What to do now

Test GLM in your own plan‑mode workflows to confirm similar performance and adjust expectations for tool calling and file generation.

Summary

GLM exhibits inconsistent clarification behavior: it under‑clarifies on ambiguous prompts (four audit tests) and over‑clarifies on degenerate inputs such as whitespace or single characters, as well as on multi‑turn answers where it re‑asks the user. Benchmark results on plan‑mode, plan‑mode_stress, tool‑calling, file_generation, and combined tasks reveal that GLM achieves 69 % success in plan‑mode, 71 % in plan‑mode_stress, 90 % in tool‑calling, 67 % in file_generation, and 75 % overall combined success. These figures are lower than those of Qwen 3‑coder‑next (92 % plan‑mode, 90 % tool‑calling) and Gemma‑4‑26b‑a4b (85 % plan‑mode, 85 % tool‑calling). The data suggest GLM’s internal ambiguity detection is unstable, leading to mixed clarification strategies.

Key changes

  • GLM under‑clarifies on ambiguous prompts (4 audit tests)
  • GLM over‑clarifies on degenerate inputs (whitespace, single char)
  • GLM over‑clarifies on multi‑turn answers (reclarify_partial_answer)
  • Plan_mode success 69% (9/13)
  • Plan_mode_stress success 71% (27/38)
  • Tool_calling success 90% (18/20)
  • File_generation success 67% (4/6)
  • Combined success 75% (58/77)

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting