🤖 AI Agent 研究Research
软件越来越多地作为科学仪器本身的一部分发挥作用,使得科学代码的失败不仅会影响程序行为,还会影响科学结论背后的证据。然而,对编码代理的现有评估在很大程度上强调了聚合任务的成功,对代理在修复科学软件时失败的原因提供了有限的见解。
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task success, providing limited insight into why agents fail when repairing scientific software.
大型语言模型代理可以通过在推理时构建工作流来适应复杂的任务,但是在一个情节中发现的程序通常在执行后被丢弃。现有技能库提供可重用的可执行例程,但通常是离线组装的,不会从客服代表自己的工作流程中增长。
Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episode are usually discarded after execution. Existing skill libraries provide reusable executable routines, but are typically assembled offline and do not grow from the agent's own workflows.
代理通过与环境的交互来学习行动,但用于培训的环境通常是围绕预定义的任务和基准手动构建或合成的。这种以任务为中心的范式使得难以扩展反映现实和不断变化的工作流程的环境,在这些工作流程中,多样化的任务可以自然地从底层世界出现。
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows where diverse tasks can naturally emerge from the underlying world.
LLM已经从语言生成器演变为能够执行复杂、长时间任务的自主代理。这种演变产生了范式,包括提示工程以引发模型功能,上下文工程以管理信息访问,利用工程来组织外部工具和资源,以及循环工程以支持持续反思和自我改进。
LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to organize external tools and resources, and Loop Engineering to support continual reflection and self-improvement.
⭐ GitHub 热门项目GitHub Trending
【GitHub】FlowOS是一个用于构建、运行、版本控制和部署AI代理和工作流的全栈平台。(⭐ 8 )
【GitHub】FlowOS is a full-stack platform for building, running, versioning, and deploying AI agents and workflows. (⭐ 8)
【GitHub】通过多智能体编排、智能路由、混合RAG、知识图表、MCP工具、持久内存和人在循环执行构建自主AI工作流程。(⭐ 5 )
【GitHub】Build autonomous AI workflows with multi-agent orchestration, intelligent routing, hybrid RAG, knowledge graphs, MCP tools, persistent memory, and human-in-the-loop execution. (⭐ 5)
【GitHub】在本地工作的开源编码代理。CLI + Web UI。支持Ollama、Claude和GPT。(⭐ 0 )
【GitHub】Open-source coding agent that works locally. CLI + Web UI. Supports Ollama, Claude, and GPT. (⭐ 0)
【GitHub】狂客同学的 Claude Code Skills 合集:开源发布/对抗性辩论/文档洁癖/公众号自动化/飞书笔记/邮件就绪检查。Kuang's daily-driver Claude Code skills, all open-sourced. (⭐ 0)
【GitHub】狂客同学的 Claude Code Skills 合集:开源发布/对抗性辩论/文档洁癖/公众号自动化/飞书笔记/邮件就绪检查。Kuang's daily-driver Claude Code skills, all open-sourced. (⭐ 0)
【GitHub】开源AI工作搜索:扫描工作门户网站,将房源评估成具有全球1-5分数的结构化A-H报告,定制您的简历,跟踪应用程序—在您的AI编码CLI中本地运行( Claude Cod (⭐ 0 )
【GitHub】Open-source AI job search: scan job portals, evaluate listings into a structured A-H report with a global 1-5 score, tailor your CV, track applications — runs locally in your AI coding CLI (Claude Cod (⭐ 0)
🚀 模型与行业动态Models & Industry
人工智能对冲基金从“华尔街的谈话”转变为“联邦传票的主题”的速度比你可以说的“多样化你的投资组合”要快。
The AI hedge fund went from "the talk of Wall Street" to "subject of federal subpoenas" faster than you can say "diversify your portfolio."
早期的测试人员对Instinct的功能赞不绝口,但有人说,人工智能助理的广泛访问权限、广泛的条款和代表用户行事的能力都带来了令人不安的权衡。
Early testers are raving about what Instinct can do, but some say the AI assistant’s sweeping access, broad terms and ability to act on users’ behalf come with uncomfortable trade-offs.
General Intuition是一家建立基础模型的初创公司,该模型训练广义人工智能代理如何在空间和时间中移动,目前正在谈判从包括Valor Ventures , Point72 Ventures在内的新投资者那里筹集60亿美元的预估值。
General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a $6 billion pre-money valuation from new investors including Valor Ventures, Point72 Ventures, a
🔥 社区热议Community
【Lobsters】热度: 23↑ | 5 评论 | 标签: vibecoding
【Lobsters】热度: 23↑ | 5 评论 | 标签: vibecoding
【HN】热度: 95 分 | 50 评论
【HN】热度: 95 分 | 50 评论