arxiv:2406.02965
Yuanhao Ban
banyh2000
AI & ML interests
None yet
Recent Activity
upvoted a paper about 14 hours ago
ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning upvoted a paper about 14 hours ago
Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards submitted a paper about 14 hours ago
Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards