论文引用、代码仓库与 GitHub Stars¶
本索引收录 132 篇在 bibliography/registry.json 中登记的核心论文与技术报告,用于提供标准 BibTeX 和官方代码状态。仓库正文引用的文献总数多于此,完整清单以各章文末的参考文献为准。论文元数据来自 Crossref、arXiv 或机构原始页面;GitHub Stars 是 2026-08-30 的快照,不代表代码质量。
- 完整 BibTeX
- 引用与仓库登记表
- Star 原始快照
- 刷新命令:
python scripts/update_bibliography.py --all - 只读一致性检查:
python scripts/update_bibliography.py --check
核心注册表边界¶
本页不是全仓逐链接书目,而是跨章节复用的核心入口。候选必须满足以下至少一项,并经人工判断为跨章节核心入口后才纳入:
- 作为机制总览、机制分章或视频基础模型主章的技术路线锚点、定义性论文或里程碑。
- 在两个及以上机制或基础模型发布页重复支撑跨路线主张,适合作为统一 BibTeX 与代码状态入口。
- 虽为机构技术报告或项目报告,但它是核心系统、模型家族或发布面的唯一一手出处。
以下内容默认留在相应章节的完整参考文献中,不机械并入核心表:
- 仅在单个任务、评测、数据集、时间线或研究记录中作局部证据的条目。
- 二手综述、新闻、聚合页,以及与核心论文无直接对应关系的社区实现。
- 仓库、许可证、标准或产品页本身;它们继续留在章末参考文献中,除非是核心系统唯一的一手报告。
本轮边界审计日期:2026-08-30。章节字母含义:
- A:经典视觉、运动与动态纹理
- B:循环预测与物理交互
- C:对抗视频生成与分布评测
- D:离散表征、token 与自回归生成
- E:扩散视频生成
- F:潜变量世界模型与控制
- G:交互式世界基础模型
- H:JEPA 表征与规划
- I:跨路线生成基础、测量与 tokenizer
- J:现代视频基础模型与系统报告
- K:因果、自回归流式生成与加速
- L:Video DiT 骨干、注意力与扩展
- M:变分随机视频与时序潜变量
- N:开放集与单/多主体视频个性化
这里的“代码仓库”表示可公开访问的 GitHub 实现;是否属于 OSI 定义的开源软件、权重是否开放,以及可否商用,仍需逐项查看仓库许可证。
仓库状态分为:官方代码(作者或机构维护)、官方研究产物(作者实验代码,但不是标准复现包)、官方相关代码(同团队的相关项目,不是该论文实现)、社区实现(第三方复现)与官方项目页(只有网页源码)。标记“未发现”的条目表示截至快照日未找到论文作者公开的 GitHub 仓库,并不等价于绝对不存在代码。
| 章节 | Cite key | 论文 / 报告 | 年份 | GitHub 与可用性 | Stars |
|---|---|---|---|---|---|
| A | doretto2003dynamic |
Dynamic Textures | 2003 | 未发现官方 GitHub 仓库 | — |
| A | horn1981determining |
Determining optical flow | 1981 | 未发现官方 GitHub 仓库 | — |
| A | schodl2000video |
Video textures | 2000 | 未发现官方 GitHub 仓库 | — |
| A | lucas1981iterative |
An Iterative Image Registration Technique with an Application to Stereo Vision | 1981 | 未发现官方 GitHub 仓库 | — |
| B | finn2016unsupervised |
Unsupervised Learning for Physical Interaction through Video Prediction | 2016 | 未发现官方 GitHub 仓库 | — |
| B | lotter2016deep |
Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning | 2016 | coxlab/prednet |
803 |
| B | mathieu2015deep |
Deep multi-scale video prediction beyond mean square error | 2015 | 未发现官方 GitHub 仓库 | — |
| B | shi2015convolutional |
Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting | 2015 | 未发现官方 GitHub 仓库 | — |
| B | srivastava2015unsupervised |
Unsupervised Learning of Video Representations using LSTMs | 2015 | 未发现官方 GitHub 仓库 | — |
| C | clark2019adversarial |
Adversarial Video Generation on Complex Datasets | 2019 | 未发现官方 GitHub 仓库 | — |
| C | tulyakov2017mocogan |
MoCoGAN: Decomposing Motion and Content for Video Generation | 2017 | sergeytulyakov/mocogan |
603 |
| C | vondrick2016generating |
Generating Videos with Scene Dynamics | 2016 | cvondrick/videogan |
706 |
| D | oord2017neural |
Neural Discrete Representation Learning | 2017 | 未发现官方 GitHub 仓库 | — |
| D | villegas2022phenaki |
Phenaki: Variable Length Video Generation From Open Domain Textual Description | 2022 | 未发现官方 GitHub 仓库 | — |
| D | yan2021videogpt |
VideoGPT: Video Generation using VQ-VAE and Transformers | 2021 | wilson1yan/VideoGPT |
1,081 |
| D | yu2022magvit |
MAGVIT: Masked Generative Video Transformer | 2022 | google-research/magvit |
1,001 |
| D | yu2023language |
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation | 2023 | google-research/magvit |
1,001 |
| E | bartal2024lumiere |
Lumiere: A Space-Time Diffusion Model for Video Generation | 2024 | 未发现官方 GitHub 仓库 | — |
| E | blattmann2023align |
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models | 2023 | 未发现官方 GitHub 仓库 | — |
| E | blattmann2023stable |
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets | 2023 | Stability-AI/generative-models |
27,276 |
| E | guo2023animatediff |
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning | 2023 | guoyww/AnimateDiff |
12,229 |
| E | ho2020denoising |
Denoising Diffusion Probabilistic Models | 2020 | hojonathanho/diffusion |
5,304 |
| E | ho2022imagen |
Imagen Video: High Definition Video Generation with Diffusion Models | 2022 | 未发现官方 GitHub 仓库 | — |
| E | ho2022video |
Video Diffusion Models | 2022 | lucidrains/video-diffusion-pytorch |
1,383 |
| E | singer2022make |
Make-A-Video: Text-to-Video Generation without Text-Video Data | 2022 | 未发现官方 GitHub 仓库 | — |
| E | openai2024sora |
Video Generation Models as World Simulators | 2024 | 未发现官方 GitHub 仓库 | — |
| F | bommasani2021opportunities |
On the Opportunities and Risks of Foundation Models | 2021 | 未发现官方 GitHub 仓库 | — |
| F | hafner2018learning |
Learning Latent Dynamics for Planning from Pixels | 2018 | google-research/planet |
1,260 |
| F | hafner2019dream |
Dream to Control: Learning Behaviors by Latent Imagination | 2019 | danijar/dreamer |
622 |
| F | hafner2023mastering |
Mastering Diverse Domains through World Models | 2023 | danijar/dreamerv3 |
3,714 |
| F | hansen2023tdmpc2 |
TD-MPC2: Scalable, Robust World Models for Continuous Control | 2023 | nicklashansen/tdmpc2 |
937 |
| F | hu2023gaia |
GAIA-1: A Generative World Model for Autonomous Driving | 2023 | 未发现官方 GitHub 仓库 | — |
| F | schrittwieser2020mastering |
Mastering Atari, Go, chess and shogi by planning with a learned model | 2020 | werner-duvaud/muzero-general |
2,863 |
| F | oh2015actionconditional |
Action-Conditional Video Prediction using Deep Networks in Atari Games | 2015 | junhyukoh/nips2015-action-conditional-video-prediction |
114 |
| F | ha2018worldmodels |
World Models | 2018 | hardmaru/WorldModelsExperiments |
734 |
| G | bruce2024genie |
Genie: Generative Interactive Environments | 2024 | 未发现官方 GitHub 仓库 | — |
| G | nvidia2025cosmos |
Cosmos World Foundation Model Platform for Physical AI | 2025 | nvidia-cosmos/cosmos-predict1 |
468 |
| G | nvidia2026cosmos3 |
Cosmos 3: Omnimodal World Models for Physical AI | 2026 | NVIDIA/Cosmos |
11,671 |
| G | valevski2024diffusion |
Diffusion Models Are Real-Time Game Engines | 2024 | GameNGen/GameNGen.github.io |
91 |
| G | yang2023interactive |
Learning Interactive Real-World Simulators | 2023 | 未发现官方 GitHub 仓库 | — |
| G | ye2026worldaction |
World Action Models are Zero-shot Policies | 2026 | dreamzero0/dreamzero |
2,603 |
| G | deepmind2024genie2 |
Genie 2: A Large-Scale Foundation World Model | 2024 | 未发现官方 GitHub 仓库 | — |
| G | deepmind2025genie3 |
Genie 3: A New Frontier for World Models | 2025 | 未发现官方 GitHub 仓库 | — |
| G | nvidia2025cosmospredict2 |
Develop Custom Physical AI Foundation Models with NVIDIA Cosmos Predict-2 | 2025 | nvidia-cosmos/cosmos-predict2 |
793 |
| G | runway2025gwm1 |
Introducing Runway GWM-1 | 2025 | 未发现官方 GitHub 仓库 | — |
| G | worldlabs2025marble |
Marble: A Multimodal World Model | 2025 | 未发现官方 GitHub 仓库 | — |
| H | assran2023selfsupervised |
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture | 2023 | facebookresearch/ijepa |
3,489 |
| H | assran2025vjepa2 |
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning | 2025 | facebookresearch/vjepa2 |
4,541 |
| H | bagatella2025tdjepa |
TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning | 2025 | facebookresearch/td_jepa |
61 |
| H | balestriero2025lejepa |
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics | 2025 | galilai-group/lejepa |
1,325 |
| H | bardes2023mcjepa |
MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features | 2023 | 未发现官方 GitHub 仓库 | — |
| H | bardes2024revisiting |
Revisiting Feature Prediction for Learning Visual Representations from Video | 2024 | facebookresearch/jepa |
4,112 |
| H | maes2026leworldmodel |
LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels | 2026 | lucas-maes/le-wm |
4,357 |
| H | murlabadia2026vjepa21 |
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning | 2026 | facebookresearch/vjepa2 |
4,541 |
| H | terver2026lightweight |
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures | 2026 | facebookresearch/eb_jepa |
765 |
| H | zhou2024dinowm |
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning | 2024 | gaoyuezhou/dino_wm |
558 |
| H | lecun2022path |
A Path Towards Autonomous Machine Intelligence | 2022 | 未发现官方 GitHub 仓库 | — |
| I | kingma2013autoencoding |
Auto-Encoding Variational Bayes | 2013 | 未发现官方 GitHub 仓库 | — |
| I | huszar2015how |
How (not) to Train your Generative Model: Scheduled Sampling, Likelihood, Adversary? | 2015 | 未发现官方 GitHub 仓库 | — |
| I | unterthiner2018towards |
Towards Accurate Generative Models of Video: A New Metric & Challenges | 2018 | 未发现官方 GitHub 仓库 | — |
| I | song2020scorebased |
Score-Based Generative Modeling through Stochastic Differential Equations | 2020 | yang-song/score_sde |
1,844 |
| I | ho2022classifierfree |
Classifier-Free Diffusion Guidance | 2022 | 未发现官方 GitHub 仓库 | — |
| I | lin2024sdxllightning |
SDXL-Lightning: Progressive Adversarial Diffusion Distillation | 2024 | 未发现官方 GitHub 仓库 | — |
| I | lin2024animatedifflightning |
AnimateDiff-Lightning: Cross-Model Diffusion Distillation | 2024 | 未发现官方 GitHub 仓库 | — |
| I | fuest2025maskflow |
MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation | 2025 | CompVis/maskflow |
28 |
| I | xie2026videorae |
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders | 2026 | 未发现官方 GitHub 仓库 | — |
| I | shutkin2026kvae |
KVAE: Family of Tokenizers for Multimodal Generative Models | 2026 | 未发现官方 GitHub 仓库 | — |
| I | guo2026vrae |
V-RAE: Rethinking Video Latent Spaces for Generation | 2026 | 未发现官方 GitHub 仓库 | — |
| J | polyak2024moviegen |
Movie Gen: A Cast of Media Foundation Models | 2024 | 未发现官方 GitHub 仓库 | — |
| J | kong2024hunyuanvideo |
HunyuanVideo: A Systematic Framework For Large Video Generative Models | 2024 | Tencent-Hunyuan/HunyuanVideo |
12,490 |
| J | hacohen2024ltxvideo |
LTX-Video: Realtime Video Latent Diffusion | 2024 | Lightricks/LTX-Video |
10,916 |
| J | ma2025stepvideo |
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model | 2025 | stepfun-ai/Step-Video-T2V |
3,188 |
| J | zheng2025opensora2 |
Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k | 2025 | hpcaitech/Open-Sora |
29,322 |
| J | wan2025wan |
Wan: Open and Advanced Large-Scale Video Generative Models | 2025 | Wan-Video/Wan2.1 |
16,908 |
| J | chen2025skyreelsv2 |
SkyReels-V2: Infinite-length Film Generative Model | 2025 | SkyworkAI/SkyReels-V2 |
7,472 |
| J | sandai2025magi1 |
MAGI-1: Autoregressive Video Generation at Scale | 2025 | SandAI-org/MAGI-1 |
3,774 |
| J | low2025ovi |
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation | 2025 | character-ai/Ovi |
1,750 |
| J | wu2025hunyuanvideo15 |
HunyuanVideo 1.5 Technical Report | 2025 | Tencent-Hunyuan/HunyuanVideo-1.5 |
4,537 |
| J | hacohen2026ltx2 |
LTX-2: Efficient Joint Audio-Visual Foundation Model | 2026 | Lightricks/LTX-2 |
9,294 |
| J | li2026skyreelsv3 |
SkyReels-V3 Technique Report | 2026 | SkyworkAI/SkyReels-V3 |
555 |
| J | seedance2026seedance2 |
Seedance 2.0: Advancing Video Generation for World Complexity | 2026 | 未发现官方 GitHub 仓库 | — |
| K | chen2024diffusionforcing |
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion | 2024 | buoyancy99/diffusion-forcing |
1,288 |
| K | yin2024causvid |
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models | 2024 | tianweiy/CausVid |
1,428 |
| K | huang2025selfforcing |
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion | 2025 | guandeh17/Self-Forcing |
3,491 |
| K | yang2025longlive |
LongLive: Real-time Interactive Long Video Generation | 2025 | NVlabs/LongLive |
2,582 |
| K | liu2025rollingforcing |
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time | 2025 | 未发现官方 GitHub 仓库 | — |
| K | yu2025videossm |
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory | 2025 | 未发现官方 GitHub 仓库 | — |
| K | samuel2026fastautoregressive |
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention | 2026 | 未发现官方 GitHub 仓库 | — |
| K | zhu2026causalforcing |
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation | 2026 | thu-ml/Causal-Forcing |
938 |
| K | xi2026quantvideogen |
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization | 2026 | 未发现官方 GitHub 仓库 | — |
| K | lv2026lightforcing |
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention | 2026 | 未发现官方 GitHub 仓库 | — |
| K | li2026rollingsink |
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion | 2026 | 未发现官方 GitHub 仓库 | — |
| K | xu2026sparseforcing |
Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation | 2026 | 未发现官方 GitHub 仓库 | — |
| K | xue2026systematic |
A Systematic Post-Train Framework for Video Generation | 2026 | 未发现官方 GitHub 仓库 | — |
| K | ji2026forcingkv |
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models | 2026 | 未发现官方 GitHub 仓库 | — |
| K | zhao2026causalforcingpp |
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation | 2026 | thu-ml/Causal-Forcing |
938 |
| K | li2026attendlocally |
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion | 2026 | 未发现官方 GitHub 仓库 | — |
| K | yesiltepe2026videomla |
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion | 2026 | 未发现官方 GitHub 仓库 | — |
| K | hu2026longliverag |
LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation | 2026 | 未发现官方 GitHub 仓库 | — |
| K | yu2026videomirai |
Video-Mirai: Autoregressive Video Diffusion Models Need Foresight | 2026 | 未发现官方 GitHub 仓库 | — |
| K | li2026aad1 |
AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation | 2026 | 未发现官方 GitHub 仓库 | — |
| K | lu2026fademem |
FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion | 2026 | 未发现官方 GitHub 仓库 | — |
| K | zheng2026causalrcm |
Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models | 2026 | NVlabs/rcm |
794 |
| K | fiebelman2026mvforcing |
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing | 2026 | 未发现官方 GitHub 仓库 | — |
| K | zhuang2026selfgradient |
Self Gradient Forcing: Native Long Video Extrapolation | 2026 | 未发现官方 GitHub 仓库 | — |
| K | xiao2026joyaivideoedit |
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion | 2026 | 未发现官方 GitHub 仓库 | — |
| K | ban2026stream4d |
Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models | 2026 | 未发现官方 GitHub 仓库 | — |
| L | peebles2022scalable |
Scalable Diffusion Models with Transformers | 2022 | facebookresearch/DiT |
8,689 |
| L | gupta2023photorealistic |
Photorealistic Video Generation with Diffusion Models | 2023 | 未发现官方 GitHub 仓库 | — |
| L | ma2024latte |
Latte: Latent Diffusion Transformer for Video Generation | 2024 | Vchitect/Latte |
1,948 |
| L | yang2024cogvideox |
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer | 2024 | zai-org/CogVideo |
12,985 |
| L | chen2025sanavideo |
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer | 2025 | NVlabs/Sana |
8,890 |
| L | huang2025linvideo |
LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation | 2025 | 未发现官方 GitHub 仓库 | — |
| L | mao2025timeripples |
Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space | 2025 | 未发现官方 GitHub 仓库 | — |
| L | chen2026sanavideo2 |
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation | 2026 | NVlabs/Sana |
8,890 |
| M | chung2015recurrent |
A Recurrent Latent Variable Model for Sequential Data | 2015 | jych/nips2015_vrnn |
291 |
| M | babaeizadeh2017stochastic |
Stochastic Variational Video Prediction | 2017 | tensorflow/tensor2tensor |
17,464 |
| M | denton2018stochastic |
Stochastic Video Generation with a Learned Prior | 2018 | edenton/svg |
188 |
| M | lee2018stochastic |
Stochastic Adversarial Video Prediction | 2018 | alexlee-gk/video_prediction |
304 |
| M | castrejon2019improved |
Improved Conditional VRNNs for Video Prediction | 2019 | facebookresearch/improved_vrnn |
39 |
| M | franceschi2020stochastic |
Stochastic Latent Residual Video Prediction | 2020 | edouardelasalles/srvp |
74 |
| M | wu2021greedy |
Greedy Hierarchical Variational Autoencoders for Large-Scale Video Prediction | 2021 | 未发现官方 GitHub 仓库 | — |
| M | saxena2021clockwork |
Clockwork Variational Autoencoders | 2021 | vaibhavsaxena11/cwvae |
50 |
| M | daniel2026latent |
Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling | 2026 | taldatech/lpwm |
133 |
| N | wang2024customvideo |
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects | 2024 | 未发现官方 GitHub 仓库 | — |
| N | ma2024magicme |
Magic-Me: Identity-Specific Video Customized Diffusion | 2024 | Zhen-Dong/Magic-Me |
458 |
| N | li2024personalvideo |
PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation | 2024 | EchoPluto/PersonalVideo |
9 |
| N | chen2025videoalchemist |
Multi-subject Open-set Personalization in Video Generation | 2025 | snap-research/MSRVTT-Personalization |
53 |
| N | liang2025movie |
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts | 2025 | 未发现官方 GitHub 仓库 | — |
| N | deng2025magref |
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement | 2025 | MAGREF-Video/MAGREF |
299 |
| N | girish2025alchemint |
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation | 2025 | snap-research/Video-AlcheMinT |
0 |
| N | huang2026rethinking |
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation | 2026 | byhuang123/PoCo |
20 |
维护约定¶
- 新增阅读条目时,先在
bibliography/registry.json登记元数据来源和仓库状态。 - 运行
python scripts/update_bibliography.py --all,同时刷新元数据、Stars、BibTeX 和本页。 - 若只想用本地快照重新生成输出,运行
python scripts/update_bibliography.py --offline。 - 官方仓库缺失时不要用社区复现冒充;如收录社区实现,必须使用
community-implementation。