{"ok":true,"article":{"slug":"the-sample-complexity-of-policy-learning-with-mu-resets-650a652f","title":"The Sample Complexity of Policy Learning with Mu-Resets","url":"https://arxiv.org/abs/2608.07772","canonical":"https://www.aimode.news/article/the-sample-complexity-of-policy-learning-with-mu-resets-650a652f","sourceName":"arXiv cs.LG","summary":"arXiv:2608.07772v1 Announce Type: new Abstract: We study policy-based reinforcement learning under the $\\mu$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $\\mu$, in addition to the starting dis…","category":"AI","image":"https://static.arxiv.org/icons/twitter/arxiv-logo-twitter-square.png","lang":"en","publishedAt":"2026-08-11T04:00:00+00:00","createdAt":"2026-08-11T19:09:30.456448+00:00"}}