Reinforcement Learning for Memory Construction in Language Models
Interested in large language models, reinforcement learning, or memory-augmented AI systems? We are seeking motivated undergraduate students to join a research project on learning what a language model should remember.
Our lab has a small language model (Qwen3-0.6B) equipped with an external 20,000-entry key–value memory that measurably improves its benchmark performance. Today that memory is constructed implicitly by gradient descent. This project treats memory construction as an explicit decision problem instead: an agent selects which entries to write, evaluates the resulting memory on held-out problems, and improves its write policy from that reward signal. The central question: can a learned write policy build a better memory than backpropagation?
Students will gain experience in:
- Large language models and external (retrieval-based) memory systems
- Reinforcement learning: multi-armed bandits and policy gradients (REINFORCE)
- Benchmark evaluation and experiment design (MMLU, GSM8K, MBPP)
- Research methodology, ablation studies, and scientific writing
Students with backgrounds in Computer Engineering, Electrical Engineering, or Computer Science are encouraged to apply. Experience with Python is required; familiarity with PyTorch or machine learning is a plus.
This is an excellent opportunity to participate in cutting-edge research at the intersection of language models and reinforcement learning, with the potential to contribute to conference publications.
Hardware-Software Codesign Lab aims to innovate at the intersection of hardware and software for energy-efficient, reliable, and secure computing systems with special foci currently at exploiting emerging memory technologies to accelerate AI and security applications. The lab projects have been supported by DARPA, NSF, semiconductor industry, etc. The lab fosters a collaborative environment with projects involving multiple graduate and undergraduate students with complementary experiences and interests.