Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation
📰 ArXiv cs.AI
Learn to extrapolate long videos using a parameter-efficient framework that preserves key elements without full generator fine-tuning
Action Steps
- Build a text-to-video diffusion transformer and freeze its weights
- Configure low-rank temporal adapters conditioned by learned shot-role prompt tokens
- Implement a recursive prompt bank to maintain long-horizon coherence
- Apply the PACR-Video framework to a multi-shot long video extrapolation task
- Test the framework's ability to preserve recurring entities, scene structure, visual style, and causal progression
Who Needs to Know This
AI researchers and engineers working on video generation tasks can benefit from this framework to improve the efficiency and coherence of their models
Key Insight
💡 Parameter-efficient frameworks can improve video extrapolation tasks without requiring full generator fine-tuning
Share This
📹 Extrapolate long videos efficiently with PACR-Video! 🤖
Key Takeaways
Learn to extrapolate long videos using a parameter-efficient framework that preserves key elements without full generator fine-tuning
Full Article
Title: Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation
Abstract:
arXiv:2607.06481v1 Announce Type: cross Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progression without full generator fine-tuning. PACR-Video keeps a text-to-video diffusion transformer frozen and augments it with low-rank temporal adapters conditioned by learned shot-role prompt tokens. To maintain long-horizon coherence, it builds a recursive prompt bank th
Abstract:
arXiv:2607.06481v1 Announce Type: cross Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progression without full generator fine-tuning. PACR-Video keeps a text-to-video diffusion transformer frozen and augments it with low-rank temporal adapters conditioned by learned shot-role prompt tokens. To maintain long-horizon coherence, it builds a recursive prompt bank th
DeepCamp AI