Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2

📰 ArXiv cs.AI

Width pruning in Llama-3.2 models improves instruction-following capabilities by up to 75%, but degrades performance on tasks relying on parametric knowledge, highlighting a dichotomy in model capabilities

advanced Published 7 May 2026
Action Steps
  1. Apply width pruning to GLU-MLP layers in Llama-3.2 using the Maximum Absolute Weight criterion
  2. Evaluate the impact of width pruning on instruction-following capabilities using IFEval metrics
  3. Compare the performance of pruned models on tasks relying on parametric knowledge, such as MMLU and GSM8K
  4. Analyze the trade-offs between instruction-following and parametric knowledge capabilities in pruned models
  5. Test the robustness of instruction-following capabilities in pruned models using various evaluation metrics
Who Needs to Know This

Researchers and developers working with large language models like Llama-3.2 can benefit from understanding the effects of width pruning on model capabilities, particularly in instruction-following tasks

Key Insight

💡 Width pruning reveals a dichotomy in Llama-3.2 model capabilities, highlighting the need for careful evaluation of trade-offs between instruction-following and parametric knowledge

Share This
💡 Width pruning in Llama-3.2 improves instruction-following by up to 75%, but degrades parametric knowledge performance

Key Takeaways

Width pruning in Llama-3.2 models improves instruction-following capabilities by up to 75%, but degrades performance on tasks relying on parametric knowledge, highlighting a dichotomy in model capabilities

Full Article

Title: Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2

Abstract:
arXiv:2512.22671v2 Announce Type: replace-cross Abstract: Structured width pruning of GLU-MLP layers, guided by the Maximum Absolute Weight (MAW) criterion, reveals a systematic dichotomy in how reducing the expansion ratio affects different model capabilities. While performance on tasks relying on parametric knowledge (e.g., MMLU, GSM8K) and perplexity metrics degrades predictably, instruction-following capabilities improve substantially (+46% to +75% in IFEval for Llama-3.2-1B and 3B models),
Read full paper → ← Back to Reads

Related Videos

5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Learn 99% of Claude in 10 Minutes (Beginner to Pro)
Learn 99% of Claude in 10 Minutes (Beginner to Pro)
AI Andy
My Custom GPT For Google Shopping Titles
My Custom GPT For Google Shopping Titles
Daryl Mander
Gemini AI + Nano Banana: Deep Research to Full eBook FAST
Gemini AI + Nano Banana: Deep Research to Full eBook FAST
LoverFighterWriter
How to Use Google Gemini AI For Beginners (Full Tutorial)
How to Use Google Gemini AI For Beginners (Full Tutorial)
LoverFighterWriter
Claude vs ChatGPT: Which AI Writer Crushes Competitors?
Claude vs ChatGPT: Which AI Writer Crushes Competitors?
LoverFighterWriter