当前位置: 首页 > news >正文

论文速读记录 | 2026.07(2)



目录
  • When Context Returns: Toward Robust Internalization in On-Policy Distillation
  • Learning to Learn with Contrastive Meta-Objective
  • Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences
  • MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration
  • Absolute Zero: Reinforced Self-play Reasoning with Zero Data
  • auto-curriculum learning (Jiang et al., 2021b)
  • Meta-Motivo(Tirinzoni 等人,2025),zero-shot goal-conditioned RL
  • Unsupervised Skill Discovery via Recurrent Skill Training
  • Learning to Discover Skills through Guidance
  • One After Another: Learning Incremental Skills for a Changing World
  • Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching
  • Horizon Generalization in Reinforcement Learning
  • HIQL: Offline Goal-Conditioned RL with Latent States as Actions
  • Contrastive Preference Learning: Learning from Human Feedback without RL
  • Few is More: Task-Efficient Skill-Discovery for Multi-Task Offline Multi-Agent Reinforcement Learning
  • Rethinking Reward Modeling in Preference-based Large Language Model Alignment
  • DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
  • Data Center Cooling System Optimization Using Offline Reinforcement Learning
  • SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
  • Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment
  • Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
  • Thinkless: LLM Learns When to Think
  • Learning to Reason without External Rewards


When Context Returns: Toward Robust Internalization in On-Policy Distillation

  • 来源:同学的最新力作。
  • arxiv:https://arxiv.org/abs/2606.11627

Learning to Learn with Contrastive Meta-Objective

  • 来源:无意中看到的,NeurIPS 2025 oral。
  • arxiv:https://arxiv.org/abs/2410.05975

(还没读。这篇文章看起来比较古典,做的是传统 ML,并不是做 llm 的。
(这个东西能用在 llm 上吗?现在看到一个东西,就会想它能否用在 llm 上

Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences

  • 来源:无意中搜到的。ICRA 2025。
  • arxiv:https://arxiv.org/abs/2409.07268
  • GitHub:https://github.com/FeiCuiLengMMbb/paper_MTPL
  • 好奇是不是 multi-type + PbRL。

MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration

  • arxiv:https://arxiv.org/abs/2006.08170
  • 来源:合作者说有趣的 skill + meta-RL 论文,ICML 2021。

Absolute Zero: Reinforced Self-play Reasoning with Zero Data

  • arxiv:https://arxiv.org/abs/2505.03335
  • 来源:neurips 2025 best paper 的一作 yue yang 的 NeurIPS 2025 spotlight 工作。被题目吸引住了,单纯好奇,想读一读。

auto-curriculum learning (Jiang et al., 2021b)

  • 来源:RSD。似乎可以做自动 curriculum learning,或许是有启发性的。

Meta-Motivo(Tirinzoni 等人,2025),zero-shot goal-conditioned RL

  • 来源:RGSD。可能包含一个技能库,也想看。速读一下就行。

Unsupervised Skill Discovery via Recurrent Skill Training

  • 来源:合作者推荐的 skill discovery 先前工作。

Learning to Discover Skills through Guidance

  • 来源:同上。

One After Another: Learning Incremental Skills for a Changing World

  • 来源:同上。

Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching

  • 来源:同上。

Horizon Generalization in Reinforcement Learning

  • arxiv:https://arxiv.org/abs/2501.02709
  • website:https://horizon-generalization.github.io/
  • 来源:Benjamin Eysenbach 的新作,是一篇 arxiv paper,同学说有趣。

HIQL: Offline Goal-Conditioned RL with Latent States as Actions

  • arxiv:https://arxiv.org/abs/2307.11949
  • website:https://seohong.me/projects/hiql/
  • 来源:合作者推荐的文章,好像也是 Benjamin Eysenbach 发表的。

Contrastive Preference Learning: Learning from Human Feedback without RL

  • arxiv:https://arxiv.org/abs/2310.13639
  • GitHub:https://github.com/jhejna/cpl
  • 来源:无意中搜到的文章,ICLR 2024,好像之前读过。
  • 主要内容:

Few is More: Task-Efficient Skill-Discovery for Multi-Task Offline Multi-Agent Reinforcement Learning

  • arxiv:https://arxiv.org/abs/2502.08985
  • 来源:同学的最新工作。
  • 主要内容:
    • 这篇文章关注的 setting 是 offline multi-task MARL;特别的,agent 只在(比如说)三个人合作的场景上训练,然后就可以泛化到任意多个人合作的场景。同学讲的故事是,用 transformer 作为一个翻译器,把三个人的合作动作翻译为多个人的,感觉这个故事听起来非常好。

Rethinking Reward Modeling in Preference-based Large Language Model Alignment

  • arxiv:https://arxiv.org/abs/2411.04991
  • OpenReview:https://openreview.net/forum?id=rfdblE10qm
  • 来源:ICLR 2025 oral。
  • 主要内容:
    • 这篇文章关注 LLM 的 RLHF。据说不采用 bradley-terry model 来建模 reward model,而是直接训一个分类器,学习一个 (x,y) 是好的还剩坏的,然后使用分类器的概率 logit 作为 RLHF 的 reward。
    • 是否使用了非成对的比较 \((x_1, y_1^+, x_2, y_2^-)\),而非把成对比较 \((x, y^+, y^-)\) 打乱(?)
    • 实验是否过于 toy(?)理论大概说了什么(?)

DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback

  • arxiv:https://arxiv.org/abs/2410.05527
  • open review:https://openreview.net/forum?id=2iYVBqRHK4
  • 来源:合作者推荐的文章。
  • 主要内容:
    • preference-based index policy(?)
  • whittle index,一个结论,两个等价条件,经典问题的证明方式。

Data Center Cooling System Optimization Using Offline Reinforcement Learning

  • arxiv:https://arxiv.org/pdf/2501.15085
  • 来源:xianyuan zhan 组的新文章。
  • 主要内容:
    • T-symmetry。

SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking

  • arxiv:https://arxiv.org/abs/2407.04752
  • 来源:师兄推荐的神秘文章,ICLR 2025 poster。

Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment

  • arxiv:https://arxiv.org/abs/2410.23680
  • 来源:偶然看到的文章。

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning

  • arxiv:https://arxiv.org/abs/2505.21067
  • 来源:偶然看到的文章。

Thinkless: LLM Learns When to Think

  • arxiv:https://arxiv.org/abs/2505.13379
  • 来源:偶然看到的文章。

Learning to Reason without External Rewards

  • arxiv:https://arxiv.org/abs/2505.19590
  • 来源:偶然看到的文章。


http://www.jsqmd.com/news/1223018/

相关文章:

  • 2026年7月最新格拉苏蒂泉州浦西万达广场维修保养服务电话 - 亨得利钟表维修中心
  • 湖南公务员考试辅导机构排行:本土与连锁实力对比 - 互联网科技品牌测评
  • 2026企业集采场景下电源线插头厂家哪家好要点解读 - 奔跑123
  • 2026年镇江企业宣传片制作广受好评服务商盘点推荐 - 奔跑123
  • 2026年保定市全城上门回收劳力士手表竞秀区向阳南大街489号赵掌柜二奢回收名表 - 优企甄选
  • 【.NET并发编程 - 20】生产环境诊断实战
  • 2026年7月最新芝柏太原吾悦广场维修保养服务电话 - 亨得利官方服务中心
  • 2026 无锡锡山区工程防水排名 TOP3|厂房、车库、小区公共区域漏水维修哪家好 - 苏易房屋修缮
  • 绥芬河赴海参崴旅游靠谱旅行社实地评测排行 - 互联网科技品牌测评
  • 承德玻璃鳞片胶泥搅拌机生产厂家选型选购实用攻略分享 - 热点品牌推荐
  • 2026 勉县上门回收黄金全实测!城乡山区全域覆盖,30 年老店在家变现零套路 - 华金汇黄金回收
  • 拒绝一口价套路!2026洛川县黄金回收避雷指南|按克实算+三十年本土老牌门店汇总 - 华金汇黄金回收
  • 2026年伽罗沉香线香品牌推荐:问菩文创天然结香 - 秋山寄远
  • 湖北虚拟资料加密分享工具指南 帮本地内容方守护数据 - 热点品牌推荐
  • 2026年7月最新福州市马尾区亨得利官方名表服务中心电话公示 - 亨得利官方博客
  • 子长全市上门回收黄金|婚嫁三金、碎金、18K 金、金条免费估价无克扣 - 华金汇黄金回收
  • 积家香港官方售后服务热线公告:2026年7月最新网点地址与客服 - 积家官方售后服务中心
  • 2026年菏泽汇图人工智能科技GEO优化服务内容一览 - 奔跑123
  • Ansys Q3D 均流分析,代理商推荐 - 品牌深度评测
  • 扬州邗江区邗上街道亨得利官方名表服务中心电话公示(2026年7月最新) - 亨得利官方
  • 2026 年现阶段通山值得关注的小区门口挡车球生产厂家有哪些,别再乱放!这小玩意儿竟是小区停车的救星 - 行业鉴选官
  • 2026连续拉伸膜包装机高质量品牌排行信息汇总 - 奔跑123
  • 2026 年 7 月求职辅导机构综合实力测评:5 家头部品牌谁是行业标杆? - 互联网科技品牌测评
  • 2026年伽罗沉香线香哪家性价比高:问菩文创值得关注 - 云溪自乐
  • 2026年菏泽GEO优化公司综合服务能力中立测评 - 奔跑123
  • 贵州木塑批发选购实用指南 各类场景高性价比货源挑选参考 - 热点品牌推荐
  • 2026 略阳黄金回收实测测评|称重验金全程公开,30 年连锁老店全域免费上门 - 华金汇黄金回收
  • 2026中国直发英国超大件庄家选择参考要点汇总 - 奔跑123
  • 2026年7月最新宇舶金华银泰百货维修保养服务电话 - 亨得利钟表维修中心
  • 2026国内口碑好的小型液压系统厂家对比评测 - 奔跑123