[NeurIPS 2025] Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
-
Updated
Nov 2, 2025 - Python
[NeurIPS 2025] Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
LLaMA-2-7B self-alignment via instruction backtranslation. 95% labeling reduction, 0.0622% trainable params (LoRA r=8), 3 HuggingFace artifacts published.
Add a description, image, and links to the self-alignment topic page so that developers can more easily learn about it.
To associate your repository with the self-alignment topic, visit your repo's landing page and select "manage topics."