Briefing

Prompt Length Matters: Qwen 3.5 vs 3.6 and Gemma 4 Performance Differences

ai-dev
by /u/Excellent_Jelly2788 ·

Run prompt tests on Qwen 3.6 and Gemma 4 to confirm that longer prompts can degrade Qwen 3.6 accuracy; adjust prompt style to match each model’s preference.

What to do now

Test your own prompts on Qwen 3.6 and Gemma 4 to determine optimal prompt length and wording for your use case.

Summary

The author ran a comparative test on three models—Qwen 3.5, Qwen 3.6, and Gemma 4—using two prompt styles: a concise version and a longer, narrative version. Each combination was executed ten times to gather statistical consistency. Results showed that most incorrect answers assumed a flat $5 per box, yielding $150 instead of the correct $300, except for Qwen 3.6 IQ2 which sometimes ignored sibling boxes. Gemma 4 performed best on the long prompt, interpreting it as a business scenario with varying buying and selling prices, while Qwen 3.6 struggled with the long prompt even in the highest quality Q8 quant, often missing the business context or sibling boxes. The IQ2 quant surprisingly delivered accurate answers in both prompt lengths. The author highlights that similar models can require different prompting styles, underscoring the importance of prompt engineering for each LLM. The full dataset and token statistics are available on the author’s website, and the comparison tool can be replicated locally. This study suggests that prompt length and wording can significantly influence model accuracy across different architectures.

Key changes

  • Qwen 3.6 accuracy drops with long prompt, often missing business context or sibling boxes.
  • Gemma 4 excels on long prompt, interpreting it as a business scenario.
  • Qwen 3.5 performs similarly to 3.6 but slightly worse on long prompt.
  • IQ2 quant consistently accurate across both prompt lengths.
  • Prompt length and wording significantly affect model accuracy.
  • Similar models may require distinct prompting strategies.
  • The study provides token counts and answer statistics for each model.

Affects

none

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting