We introduce DADiff, a diffusion-based framework for online dynamics adaptation in reinforcement learning that measures dynamics mismatch through generative trajectory deviation.
We present the first systematic study of multilingual instruction following in vision-language-action models, identify the causes of the multilingual gap, and propose MPCA to align multilingual embeddings.
Welcome to my personal blog! I will share research insights, learning notes, and thoughts on reinforcement learning, data mining, and large language models.