Research Showcase
Online Reflection for Self-Improving Language Models
A novel framework that converts cross-episode signals into reusable procedural memory and distills it into the policy's weights during training. Memory functions as a training scaffold, enabling memory-free inference while achieving significant improvements over existing methods.
Learning from Language Feedback via a Co-Evolutionary EM Framework
VPD reframes on-policy self-distillation as a Variational Expectation-Maximization problem. The teacher is actively refined on trajectory outcomes (E-step) and then distilled into the student's own rollouts (M-step), with a dynamic trust region anchored to the current policy โ co-evolving teacher and student inside a single shared-weight network.
Stay tuned for more groundbreaking research in artificial intelligence, machine learning, and natural language processing.
Developing AI systems that learn from their own experiences and continuously improve their capabilities over time.
Exploring how AI can effectively store, retrieve, and utilize knowledge to enhance reasoning and decision-making.
Advancing RL techniques to enable more efficient and effective learning from verifiable rewards and feedback.