🤖 AI Agent 研究Research
物理理解和推理依赖于形成世界的紧凑和可概括的表示。虽然现代视觉语言模型可以识别和解释不同的物理事件,但它们通常缺乏对潜在机制的明确表示,例如对象状态、物理参数和管理动态,这对于可靠地推理世界如何演变和对干预措施的反应是必需的。
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and governing dynamics-needed for reliably reasoning how the world evolves and responds to interventions.
长期代理任务需要大型语言模型( LLM )来迭代检索、整合和维护多轮交互中的分散信息,但保留所有交互历史会导致工作环境不断增长。为了弥合这些差距,我们引入了ContextPilot ,这是一种用于长期代理推理的主动上下文管理框架。
Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning.
长时程 Agent 运行产生的经验可以改善当前的运行和未来的工作。大多数自我改进方法仅在执行结束后处理此体验,因此它们无法重定向正在进行的运行,也无法立即应用和验证从中学到的经验教训。我们认为,自我改进应该是实时的,利用不断涌现的经验来重定向当前运行并更新持久化 harness。
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness.
缩放世界模型的一个常见策略是使用更多的计算来训练更多的抓取视频。我们认为这种策略效率低下:扩展世界模型还需要一个提供接地奖励信号的递归数据引擎。由于代码是可执行的,编译器和运行时可以为LLM的强化学习( RL )后培训提供高质量的奖励。
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs.
⭐ GitHub 热门项目GitHub Trending
Habitus-AI是AI Agent记忆系统和认知 harness的根本新方法。双Cypher系统沿着垂直图运行。实时突变边缘。培养习惯和动态人格特质。
Habitus-AI is fundamentally new approach to AI Agent Memory Systems and Cognitive Harness. Dual-Cypher system runs along perpendicular graphs. Mutate edges in real time. Builds habits and dynamic personality traits.
用于编码和使用AI代理工具的开源代理执行恢复控制器。
Open-source agent execution-recovery controller for coding and tool-using AI agents.
Claude Code、Cowork、Codex和其他AI代理的AI产品管理技能和插件:证据标记的PRD、规格、要求、大米优先级、积压和路线图评分、产品战略、GTM发布计划、发布验证、基准测试包、Web和移动应用程序的UX/UI设计提示。
AI product management skills and plugin for Claude Code, Cowork, Codex and other AI agents: evidence-tagged PRDs, specs, requirements, RICE prioritization, backlog and roadmap scoring, product strategy, GTM launch plans, release verification, benchmark packs, UX/UI design prompts for web and mobile apps.
交出工作。取回结果。您的团队的开源人工智能同事:拥有自己云计算机的代理、您的工具和上下文,交还已完成的工作-网站、幻灯片、电子表格、报告、公关。在您的订阅上运行Claude Code、Codex、OpenCode。
Hand off the work. Get back the result. The open-source AI coworker for your team: agents with their own cloud computer, your tools and context, handing back finished work - websites, decks, spreadsheets, reports, PRs. Runs Claude Code, Codex, OpenCode on your subscription.
一种语义、事务、可重播的MCP控制平面,让自主人工智能代理将Dwarf Fortress作为一个长期存在的文明进行操作。Safe-Rust phase-0B脚手架: MVCC世界状态,目睹语义计划,证据支持的效果,确定性重放。
A semantic, transactional, replayable MCP control plane that lets autonomous AI agents operate Dwarf Fortress as a long-lived civilization. Safe-Rust phase-0B scaffold: MVCC world state, witnessed semantic plans, evidence-backed effects, deterministic replay.
🚀 模型与行业动态Models & Industry
卡特彼勒花费了数十年的时间,让自动机器在偏远的采矿地点工作。现在,它正在将这种体验带入人工智能部署。
Caterpillar has spent decades putting autonomous machines to work at remote mining sites. It's now bringing that experience to AI deployment.
针对特定未对齐行为的10个基准,自动化系统能够提高每个基准的性能,而不会降低整体性能。
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.
这些更新表明,谷歌希望将AI Mode定位为某种程度上的AI旅行社,因为它不仅仅是帮助用户查找信息,而是实际处理行程规划和预订流程的一部分。
The updates indicate that Google is looking to position AI Mode as an AI travel agent of sorts, as it's moving beyond simply helping users find information to actually handling parts of the trip-planning and booking process.
🔥 社区热议Community
【Lobsters】热度: 100↑ | 20 评论 | 标签: philosophy, vibecoding
【Lobsters】热度: 100↑ | 20 评论 | 标签: philosophy, vibecoding
【Lobsters】热度: 18↑ | 20 评论 | 标签: vibecoding
【Lobsters】热度: 18↑ | 20 评论 | 标签: vibecoding