Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction

📰 ArXiv cs.AI

Spectral Compact Training (SCT) is a method for pre-training large language models using permanent truncated SVD and Stiefel QR retraction to reduce memory usage

advanced Published 2 Apr 2026
Action Steps
  1. Replace dense weight matrices with permanent truncated SVD factors
  2. Use standard backpropagation to flow gradients through compact spectral factors
  3. Retract U, V to the Stiefel manifold to maintain orthogonality
  4. Train large language models using the compact spectral factors without materializing the full dense matrix
Who Needs to Know This

ML researchers and engineers working on large language models can benefit from this method to reduce memory usage and improve training efficiency. This can be particularly useful for teams working on consumer hardware with limited memory resources

Key Insight

💡 SCT enables efficient pre-training of large language models on consumer hardware by reducing memory usage

Share This
🔍 Reduce memory usage for large language models with Spectral Compact Training (SCT) 🚀

Key Takeaways

Spectral Compact Training (SCT) is a method for pre-training large language models using permanent truncated SVD and Stiefel QR retraction to reduce memory usage

Full Article

Title: Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction

Abstract:
arXiv:2604.00733v1 Announce Type: cross Abstract: The memory wall remains the primary bottleneck for training large language models on consumer hardware. We introduce Spectral Compact Training (SCT), a method that replaces dense weight matrices with permanent truncated SVD factors W = U diag(s) V^T, where the full dense matrix is never materialized during training or inference. Gradients flow through the compact spectral factors via standard backpropagation, and U, V are retracted to the Stiefel
Read full paper → ← Back to Reads

Related Videos

5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Gemini AI + Nano Banana: Deep Research to Full eBook FAST
Gemini AI + Nano Banana: Deep Research to Full eBook FAST
LoverFighterWriter
How to Use Google Gemini AI For Beginners (Full Tutorial)
How to Use Google Gemini AI For Beginners (Full Tutorial)
LoverFighterWriter
Claude vs ChatGPT: Which AI Writer Crushes Competitors?
Claude vs ChatGPT: Which AI Writer Crushes Competitors?
LoverFighterWriter
Off-Page Topical Map: Why Third-Party Corroboration Improves LLM Visibility (Karl ft James)
Off-Page Topical Map: Why Third-Party Corroboration Improves LLM Visibility (Karl ft James)
James Dooley
AI Reputation Tree - Getting The LLMs To Be Your 24/7 Sales Engine (Karl Hudson ft James Dooley)
AI Reputation Tree - Getting The LLMs To Be Your 24/7 Sales Engine (Karl Hudson ft James Dooley)
James Dooley