二、GitHub 当日热点:Agent 与机器人方向的热门仓库与方法
Part 2 · GitHub: trending agent and robotics repositories
本期仓库集中在「协作与记忆的工程实现」:Git 原生记忆、意图冲突协议、人类与 Agent 共用画布、以及统一 CLI 管理异种机器人。
These repos cluster around engineering collaboration and memory: Git-native memory, an intent-conflict protocol, a shared human-agent canvas, and one CLI for heterogeneous robots.
01
KKKKhazix/AIHOT — 自动找热点、自动写日报的站点框架
KKKKhazix/AIHOT — a framework that finds trends and writes the digest itself
⭐ 5,947 · TypeScript · 2026-09-28 创建 · 2026-10-04 更新⭐ 5,947 · TypeScript · created 2026-09-28 · pushed 2026-10-04
AIHOT 是一个自动抓取热点并生成日报的网站框架:把信源与精选标准换成你自己的,它就变成你的行业热点站[23]。它面向的是想长期运营垂直资讯站、又不愿每天手工挑内容的个人与团队。实现上的关键点是信源抽象与精选标准可配置——聚合、打分、成稿三个环节解耦,因此可以只替换其中一层(例如换掉打分规则)。值得借鉴的是把「编辑标准」写成配置而不是写成提示词;风险是自动化日报的信噪比完全取决于信源质量与打分规则,接入前应先用历史数据回测命中率。
AIHOT is a framework that crawls trending topics and generates a daily digest — swap the sources and curation criteria and it becomes your own industry news site[23]. It targets individuals and teams who want to run a vertical news property without hand-picking content daily. The key design is that source abstraction and curation criteria are configurable, decoupling aggregation, scoring and drafting, so you can replace just one layer, such as the scoring rule. Worth borrowing is writing the editorial standard as configuration rather than prompts; the risk is that digest signal-to-noise depends entirely on sources and scoring, so back-test the hit rate before adopting.
02
rehan-remade/universal-modder — 让 Claude Code 去改任何 PC 游戏
rehan-remade/universal-modder — pointing Claude Code at any PC game
⭐ 3,604 · Python · 2026-09-30 创建 · 2026-10-05 更新⭐ 3,604 · Python · created 2026-09-30 · pushed 2026-10-05
这个项目的做法是把 Claude Code 指向任意 PC 游戏,通过技能、工具与 fal 的 MCP 服务完成侦察、逆向、素材生成、游戏内测试到展示视频的全流程 mod 制作[24]。它解决的问题很具体:游戏 mod 的门槛分散在逆向工程、资源打包与美术制作上,单人很难全流程走通。实现上把每个环节做成可调用的技能与工具,并用生成式模型补齐美术、3D 与音频。值得借鉴的是把长流程拆成可独立验证的环节,并在每个环节安排真实执行反馈(游戏内测试);风险是与版权、反作弊和厂商条款的冲突,发布 mod 前需自行评估法律边界。
This project points Claude Code at any PC game and uses skills, tools and fal's MCP service to cover recon, reverse engineering, generated art/3D/audio, in-game testing and showcase videos[24]. The problem is concrete: game modding spans reverse engineering, asset packing and art, which is hard for one person to cover end to end. Every stage is exposed as an invocable skill or tool, with generative models filling in art, 3D and audio. Worth borrowing is splitting a long pipeline into independently verifiable stages with real execution feedback (in-game testing); the risk is conflict with copyright, anti-cheat and vendor terms, so assess legal boundaries before shipping a mod.
03
ZJU-REAL/Easel — 覆盖发现、创作与分发的社交媒体 Agent
ZJU-REAL/Easel — a social media agent covering discovery, creation and publishing
⭐ 3,094 · Python · 2026-08-28 创建 · 2026-10-05 更新⭐ 3,094 · Python · created 2026-08-28 · pushed 2026-10-05
Easel 是一个开源社交媒体 Agent:发现热点趋势、创作内容、一键发布到小红书/抖音/知乎/B 站等平台,并学习分析哪些内容真正有效[25]。它要解决的是多平台运营的重复劳动与「发完不知道为什么有效」的问题:创作与复盘长期脱节。实现上把趋势发现、内容生成、发布与效果分析串成闭环,并用 MCP 对接平台能力。值得借鉴的是把「效果回流」放进 Agent 的学习回路,而不是停在发布成功;风险是平台条款对自动化发布通常有限制,账号安全与合规需要自行评估。
Easel is an open-source social media agent: discover trends, create content, publish everywhere across Xiaohongshu, Douyin, Zhihu and Bilibili, and learn what actually works[25]. It addresses duplicated effort across platforms and the “published but no idea why it worked” gap, where creation and review stay disconnected. Trend discovery, generation, publishing and performance analysis are chained into a loop, with MCP used to reach platform capabilities. Worth borrowing is putting performance feedback into the agent's learning loop rather than stopping at a successful publish; the risk is that platform terms usually restrict automated posting.
04
feder-cr/dots — 自带浏览器、且不容易被封的网页 Agent
feder-cr/dots — a web agent with its own browser that does not get blocked
⭐ 2,610 · Python · 2026-09-29 创建 · 2026-10-03 更新⭐ 2,610 · Python · created 2026-09-29 · pushed 2026-10-03
dots 的定位是网页 Agent 的开放实现:给它一个自己的浏览器,并且不容易被反自动化机制拦截[26]。它要解决的是网页 Agent 最常见的失败模式——被检测为机器人后任务中断,而不是推理出错。实现上把浏览器控制、反检测配置与任务执行分层,让 Agent 在接近真实用户的环境里操作。值得借鉴的是把「不被封」当成与推理能力同等重要的工程指标;风险是这类反检测能力天然与平台条款冲突,用于第三方站点抓取或高频操作时存在法律与封号风险,需谨慎评估使用场景。
dots positions itself as an open implementation of a web agent: one with its own browser that does not get blocked[26]. It attacks the most common web-agent failure, being detected as a bot and cut off rather than reasoning badly. Browser control, anti-detection configuration and task execution are layered so the agent operates in an environment close to a real user. Worth borrowing is treating “does not get blocked” as an engineering metric on par with reasoning quality; the risk is that such anti-detection capability inherently conflicts with platform terms, carrying legal and ban exposure on third-party sites.
05
kgoedecke/doop — 人类与 Agent 同场协作的设计画布
kgoedecke/doop — a design canvas where humans and agents work together live
⭐ 810 · TypeScript · 2026-08-22 创建 · 2026-10-05 更新⭐ 810 · TypeScript · created 2026-08-22 · pushed 2026-10-05
doop 是面向设计协作的多人在线画布:人类与 AI Agent 在同一个实时画布上一起设计,并内置 MCP 支持[27]。它要解决的是设计流程里人与 Agent 的工作产物割裂:Agent 生成的稿子要手动搬进设计工具,上下文随之丢失。做法上让双方共享同一份画布状态,Agent 的产出直接落在协作空间里。值得借鉴的是把 Agent 当成协作空间的参与者而不是外部生成器;风险是多人在线与模型权限叠加后,冲突合并与操作审计的复杂度显著上升,团队需先定义谁拥有最终修改权。
doop is a multiplayer canvas for design collaboration where humans and AI agents design together live, with MCP built in[27]. It tackles the split between human and agent artefacts in design workflows: agent output must be manually dragged into the design tool and context is lost. Both sides share one canvas state so agent output lands directly in the shared space. Worth borrowing is treating the agent as a participant in the collaboration space rather than an external generator; the risk is that multiplayer plus model permissions makes conflict resolution and action auditing much harder.
06
okf-memory/okf-agent-memory — Git 原生的编码 Agent 持久记忆
okf-memory/okf-agent-memory — Git-native persistent memory for coding agents
⭐ 753 · Go · 2026-09-05 创建 · 2026-10-04 更新⭐ 753 · Go · created 2026-09-05 · pushed 2026-10-04
这个项目为编码 Agent 提供Git 原生的持久记忆:实现 Google OKF v0.2,内存内 BM25 检索低于 300 微秒,内嵌 MCP 服务并支持渐进披露(progressive disclosure),据称把 token 膨胀削减约 80%,且不依赖任何外部数据库[28]。它要解决的是记忆系统的运维负担:为了持久记忆引入向量库与索引服务,本身就成了新故障源。做法上把记忆放在 Git 管理的文件里并用内存索引,使记忆可版本化、可审计、可回滚。值得借鉴的是用 Git 承担版本与审计职责,把系统复杂度留在检索层;风险是「约 80% token 削减」等数字来自项目自述,需要独立复现,且大仓库下的索引成本尚未公开。
This project gives coding agents Git-native persistent memory: Google OKF v0.2, in-memory BM25 search under 300µs, an embedded MCP server and progressive disclosure, claiming roughly 80% less token bloat with no external databases[28]. It targets the operational burden of memory systems, where adding a vector store and index service becomes a new failure source. Memory lives in Git-managed files with an in-memory index, so it can be versioned, audited and rolled back. Worth borrowing is letting Git own versioning and audit while keeping complexity in the retrieval layer; the risk is that the 80% figure is self-reported and needs independent reproduction, with index cost at large repo sizes undisclosed.
07
naw103/foremerge — 在代码冲突之前拦截意图冲突
naw103/foremerge — catching intent conflicts before code conflicts
⭐ 524 · Rust · 2026-08-21 创建 · 2026-10-04 更新⭐ 524 · Rust · created 2026-08-21 · pushed 2026-10-04
foremerge 是一个构建在 Git 之上的编码 Agent 协作协议,目标是在代码冲突发生之前先发现意图冲突[29]。它要解决的是多 Agent 并行开发的新问题:两个 Agent 各自都能通过测试,但设计目标互相矛盾,等到合并时才发现,返工成本极高。做法上把「意图」显式记录下来,并在提交前做一致性检查。值得借鉴的是把冲突检测从文本层前移到意图层,这对并行 Agent 尤其关键;风险是意图描述本身可能不准确或不完整,检查效果取决于 Agent 是否如实声明目标,协议采纳程度也仍需观察。
foremerge is a coordination protocol for coding agents built above Git that aims to catch intent conflicts before code conflicts[29]. It addresses a new problem in parallel agent development: two agents each pass their tests while their design goals contradict each other, discovered only at merge time at high rework cost. Intent is recorded explicitly and checked for consistency before commit. Worth borrowing is moving conflict detection from the text layer up to the intent layer, which matters especially for parallel agents; the risk is that intent statements may be inaccurate or incomplete, so the check is only as good as the agents' honesty.
08
showlab/Show-Harness — 一个 VLM Agent 就能玩机器人
showlab/Show-Harness — just a VLM agent can play robots
⭐ 509 · Python · 2026-09-07 创建 · 2026-10-03 更新⭐ 509 · Python · created 2026-09-07 · pushed 2026-10-03
Show-Harness 的主张很短:「一个 VLM Agent 就能操作机器人」,即用视觉语言模型作为控制栈的核心,配合 harness 把感知与动作组织起来[30]。它要解决的是机器人软件栈过重的问题:传统方案需要专门的控制策略、状态机与大量手工工程。做法上把通用 VLM 当作决策层,用 harness 补上执行与反馈。值得借鉴的是优先验证通用模型加薄 harness 能否覆盖场景,再决定是否投入专用策略;风险是这类方案在精度要求高或安全关键的接触任务上通常不足,且真实机器人上的鲁棒性需要独立验证。
Show-Harness makes a short claim: just a VLM agent can play robots — using a vision-language model as the core of the control stack, with a harness organising perception and action[30]. It addresses an over-heavy robotics stack where traditional approaches need bespoke control policies, state machines and much manual engineering. A general VLM serves as the decision layer while the harness supplies execution and feedback. Worth borrowing is validating whether a general model plus a thin harness covers the scenario before investing in specialist policies; the risk is that such approaches usually fall short on high-precision or safety-critical contact tasks.
09
rokbenko/quackd — 一条命令管理所有机器人
rokbenko/quackd — one CLI for all your robots
⭐ 250 · Python · 2026-08-28 创建 · 2026-10-03 更新⭐ 250 · Python · created 2026-08-28 · pushed 2026-10-03
quackd 提供一条统一 CLI 连接、下发指令并协调多台机器人,每台以 LLM 作为大脑(Claude、OpenAI、Gemini、Grok 或通过 Ollama/vLLM 本地部署),并可用 VLA 驱动机械臂与决策模型,已适配 Microduck、Open Duck Mini、LeRobot、AlohaMini、ToddlerBot 与任意 ROS 2 底盘[31]。它要解决的是机器人生态碎片化:每家硬件一套 SDK,很难混用。做法上把硬件抽象成统一接口,把智能放在外置主机。值得借鉴的是硬件抽象层与模型层解耦,使更换模型或机器人不需要重写整套流程;风险是抽象层会掩盖硬件特有约束,安全关键动作仍需逐机型验证。
quackd offers one CLI to connect, command and coordinate multiple robots, each with an LLM as its brain (Claude, OpenAI, Gemini, Grok, or local via Ollama/vLLM) and VLAs driving arms and decision models, supporting Microduck, Open Duck Mini, LeRobot, AlohaMini, ToddlerBot and any ROS 2 base[31]. It addresses robotics ecosystem fragmentation, where each vendor ships its own SDK. Hardware is abstracted behind one interface while intelligence runs offboard. Worth borrowing is decoupling the hardware abstraction layer from the model layer; the risk is that abstraction hides hardware-specific constraints, so safety-critical motions still need per-model validation.
10
AskTheWay/dsh-auto-memory — 自动记忆插件:把 MEMORY.md 注入系统提示
AskTheWay/dsh-auto-memory — auto-injecting MEMORY.md into the system prompt
⭐ 88 · TypeScript · 2026-09-22 创建 · 2026-10-05 更新⭐ 88 · TypeScript · created 2026-09-22 · pushed 2026-10-05
这个插件为 DeepSeek Harness 提供Claude Code 风格的自动记忆:类型化记忆文件加 MEMORY.md 索引,自动注入系统提示,纯文件实现、不依赖外部服务[32]。它要解决的是记忆系统的运维成本:为持久记忆部署数据库与检索服务,本身就成为新的故障点。做法上用文件承载记忆、用索引文件控制注入内容,把复杂度压到最低。值得借鉴的是「文件 + 索引 + 自动注入」这套最小可行记忆方案,它足够透明也容易回滚;风险是知识量增长后索引本身会膨胀,如何筛选注入内容将成为新的瓶颈。
This plugin gives DeepSeek Harness Claude Code–style auto-memory: typed memory files plus a MEMORY.md index auto-injected into the system prompt, file-only with no external services[32]. It targets the operational cost of memory systems, where deploying a database and retrieval service for persistent memory becomes a new failure point. Files carry memory and an index file controls what gets injected, keeping complexity minimal. Worth borrowing is the minimal viable memory recipe of files plus index plus auto-injection, which is transparent and easy to roll back; the risk is that the index itself grows with knowledge, making selection the next bottleneck.
三、每日论文:arXiv 上的 Agent 研究
Part 3 · Daily Papers: agent research on arXiv
本期 8 篇覆盖 harness 自进化、终端 Agent 的信用分配、浏览 Agent 的多语言压力测试、安全评测效度与成本感知的进化搜索。
These eight papers cover harness self-improvement, credit assignment for terminal agents, a multilingual browsing stress test, the validity of security benchmarks and cost-aware evolutionary search.
01
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
2610.03548 · cs.AI · 2026-10-022610.03548 · cs.AI · 2026-10-02
这篇论文提出任务与 harness 的共同进化:现有递归式推理数据合成只把生成出的题目当作新种子复用,却从不修改构造题目的 harness 本身,因此难度提升很快碰到天花板[33]。它要解决的问题是推理数据合成无法自我升级。方法是让 harness 在线自我改进——把求解器中途的失败转成可复用技能,并在每批任务结束后修订技能、提示词与工作流,只有在新方案能在成本上限内产出更难且有效的任务时才采纳;模型权重与验证标准保持固定,保证改进来自 harness 而不是模型或评分放水。对做合成数据与自进化 Agent 的团队,这条直接可借鉴的是把「生成器」也纳入优化对象,并给定成本预算作为准入条件;限制是效果依赖验证标准的可靠性,评分一旦松动改进就失去意义。
This paper proposes task–harness co-evolution: existing recursive reasoning-data synthesis reuses generated problems as new seeds but never changes the harness that constructs them, so difficulty quickly plateaus[33]. The problem is that synthesis cannot upgrade itself. The method lets the harness self-improve online — intermediate solver failures become reusable skills, and after each batch skills, prompts and workflows are revised, with candidates adopted only if they yield harder valid tasks within a bounded cost increase; model weights and verification criteria stay fixed, so gains come from the harness rather than the model or looser grading. For synthesis and self-evolving agent teams the transferable idea is treating the generator as an optimisation target with a cost budget as the admission test; the limit is dependence on verification integrity.
02
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
2610.03634 · cs.AI · 2026-10-022610.03634 · cs.AI · 2026-10-02
DepGPO 针对终端 Agent 的信用分配问题:在编码、调试这类多步终端任务里,后面的命令往往依赖前面命令产生的结果,但现有轨迹级与步骤级信用分配都不会显式追踪「读—写依赖」,于是训练信号被分给了无关操作[34]。这直接削弱了从真正关键步骤学习的能力。方法上用命令之间的执行依赖来指导信用分配,把功劳与责任沿着依赖链传递。对训练终端或编码 Agent 的团队,值得借鉴的是把执行依赖图当作训练信号的一部分,而不是只看最终成败的平均回报;限制是依赖解析本身可能出错,复杂 shell 管道与副作用的静态分析仍是难点。
DepGPO targets credit assignment in terminal agents: in multi-step coding and debugging tasks later commands depend on results produced earlier, yet existing trajectory-level and step-level methods never trace those read-write dependencies, so training signal goes to irrelevant operations[34]. That weakens learning from the steps that actually mattered. The method uses execution dependencies between commands to guide credit assignment, propagating credit along the dependency chain. For teams training terminal or coding agents the transferable idea is using the execution dependency graph as part of the training signal instead of mean terminal reward; the limit is that dependency parsing can itself be wrong, especially with complex shell pipelines and side effects.
03
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
2610.03574 · cs.AI, cs.LG · 2026-10-022610.03574 · cs.AI, cs.LG · 2026-10-02
HyperBrowseComp 是一个多语言、多模态的浏览 Agent 压力测试集:423 道人工编写并人工校验的题目,覆盖 13 种语言,由母语或高熟练度作者撰写[35]。它针对的是现有浏览基准偏简单、偏英语、偏文本的问题:真正难的检索需要定位冷门证据、跟随多步线索链,或检查视频、扫描件、图片与地图等异构来源。设计上还有一道关键过滤:先用无联网模型筛掉简单题,降低仅靠参数记忆就能作答的概率。对做浏览器 Agent 评测的团队,值得借鉴的是用「无联网模型能否答对」作为题目难度的准入门槛,同时把跨语言与跨模态纳入同一套评分;限制是 423 道的规模仍偏小,且答案虽可公开验证,难度分布尚未充分公开。
HyperBrowseComp is a multilingual, multimodal browsing stress test: 423 hand-authored, human-validated questions across 13 languages written by native or highly proficient speakers[35]. It targets the easy, English-centric, text-heavy bias of existing browsing benchmarks: genuinely hard retrieval requires locating obscure evidence, following multi-step clue chains, or inspecting videos, scans, images and maps. A key filter is that easy questions are removed using models without internet access, reducing the chance that parametric memory alone suffices. For browse-agent evaluation the transferable idea is using offline models' ability to answer as the difficulty admission test; the limit is that 423 questions is still small.
04
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
2610.03585 · cs.CR, cs.AI, cs.LG · 2026-10-022610.03585 · cs.CR, cs.AI, cs.LG · 2026-10-02
这篇论文质疑 Agent 安全评测的效度:安全基准常用攻击成功率(ASR)来衡量鲁棒性,并据此比较模型与防御方案,默认这个分数描述了 Agent 的安全性[36]。作者提出威胁保持的表示敏感性(TPRS)来检验这个假设——在任务、有害动作、安全策略、真值、环境与评测标准都不变的前提下,只改变 Agent 可见的表示形式,看 ASR 变化多少。结论是在 Agent Security Bench 上,仅表示层的改动就显著改变了 ASR,说明分数在很大程度上测的是「表述」而不是「威胁」。对做 Agent 红队与安全评测的团队,参考价值是评测设计本身必须先做敏感性检验,否则防守方可能只是在优化题型;限制是这类检验需要额外工程,且论文结论建立在单一基准上。
This paper questions the validity of agent security evaluation: benchmarks use attack success rate (ASR) as the measure of robustness and compare models and defences on it, assuming the score describes agent security[36]. The authors introduce threat-preserving representation sensitivity (TPRS): holding task, harmful action, security policy, ground truth, environment and criteria fixed, change only the agent-visible representation and observe how far ASR moves. On Agent Security Bench, representation changes alone shift ASR substantially, meaning the score largely measures phrasing rather than threat. For red-teaming and security evaluation teams the reference is that the evaluation design itself needs a sensitivity check, or defenders optimise for question format; the limit is the extra engineering and a single benchmark basis.
05
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
2610.03675 · cs.NE, cs.AI, cs.CL · 2026-10-022610.03675 · cs.NE, cs.AI, cs.CL · 2026-10-02
FrugalEvo 指出 LLM 引导的程序进化(如 AlphaEvolve 一类方法)都在固定迭代次数下优化性能增益,而忽略了成本,但工程上真正要最大化的是「每单位成本带来的增益」[37]。方法上做角色分工:用更强也更贵的 LLM 探索解题策略,用更便宜的 LLM 负责实现并迭代精修代码,从而把昂贵的探索与廉价的实现分开。此外设计了缓存友好的进化流程,通过 harness 与提示词让不同进化步骤最大化共享前缀,从而提升缓存命中、压低成本。对做自动化优化与 Agent 长期任务的团队,值得借鉴的是把「成本感知」写进搜索目标,并用模型分层来匹配任务难度;限制是分层策略需要事先知道各类子任务所需的模型档位,判断错了会牺牲增益。
FrugalEvo observes that LLM-guided program evolution (AlphaEvolve-style methods) optimises performance gain over a fixed number of iterations while ignoring cost, whereas engineering wants to maximise gain per unit cost[37]. It splits roles: a stronger, costlier LLM explores solution strategies while a cheaper LLM implements and iteratively refines the code, separating expensive exploration from cheap implementation. It also designs a cache-efficient evolution process where the harness and prompts maximise shared prefixes across steps to raise cache hits and cut cost. For automated optimisation and long-horizon agent teams the transferable idea is writing cost-awareness into the search objective and matching model tier to task difficulty; the limit is that the tiering must be chosen correctly up front.
06
What Should World Models Forget? Stratified Retention for Continual Adaptation
What Should World Models Forget? Stratified Retention for Continual Adaptation
2610.03713 · cs.LG, cs.AI, cs.CV · 2026-10-022610.03713 · cs.LG, cs.AI, cs.CV · 2026-10-02
这篇论文挑战了持续学习的一条默认规则:它把「旧数据上性能下降」当作失败,这个惯例继承自预测目标稳定的场景——在那里正确标签永远正确[38]。世界模型不满足这个条件:它的预测目标是会变化的环境,因此曾经正确的知识后来可能变错,丢弃它是必要行为而不是缺陷。作者把这个问题形式化为世界模型特有的非平稳真值,并指出世界模型的独特之处在于同时还编码了永远不该被修正的知识,因此需要分层保留(stratified retention)而非统一抗遗忘。对做世界模型与长期记忆的团队,值得借鉴的是给知识分层:可过期、需保留、永不可改,并分别设计更新策略;限制是分层标准如何自动判定尚未解决,仍需人工或额外监督。
This paper challenges a default rule of continual learning: treating degradation on previously seen data as failure, a convention inherited from settings with a stationary prediction target where a correct label stays correct forever[38]. World models do not satisfy that: their target is a changing environment, so knowledge once accurate can become false, and discarding it is required behaviour rather than a defect. The authors formalise this non-stationary ground truth and note the distinctive part: world models also encode knowledge that must never be revised, hence stratified retention instead of uniform anti-forgetting. For world-model and long-term memory teams the transferable idea is stratifying knowledge into expirable, retainable and immutable, each with its own update policy; the limit is how to decide the strata automatically.
07
Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows
Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows
2610.03564 · cs.AI · 2026-10-022610.03564 · cs.AI · 2026-10-02
这篇论文用金融 Agent 做了一次很干净的消融:FinSkillBench 覆盖组合构建、风险管理与基本面分析三个方向的 12 个子任务、2,603 个时点片段,配有可再生的隐藏真值与任务专属的确定性验证器;在 9 个模型、3 种资源条件下跑了 17,820 个片段[39]。关键结论是:人工整理的技能包把平均分从 0.366 提升到 0.528(+16.2 分),而在单次任务内临时生成的技能只带来 +0.5 分。也就是说,Agent 在可验证工作流里的增益主要来自「外部整理好的过程性资源」,而不是模型临场发挥。对做金融或任何可验证工作流的团队,值得借鉴的是把过程性知识产品化(技能包),并建立确定性验证器;限制是结论依赖具体任务族,迁移到开放式任务时需要重新验证。
This paper runs a clean ablation on financial agents: FinSkillBench covers 2,603 point-in-time episodes across 12 subtasks in portfolio construction, risk management and fundamental analysis, with hidden regenerable ground truth and deterministic task-specific verifiers; 17,820 episodes were executed across 9 models and 3 resource conditions[39]. The headline result: curated skill packages raise mean scores from 0.366 to 0.528 (+16.2 points), while skills generated within a single episode add only +0.5 points. In verifiable workflows the gain comes from externally curated procedural resources, not on-the-fly model improvisation. For financial and other verifiable workflows the transferable idea is productising procedural knowledge as skill packages and building deterministic verifiers; the limit is dependence on these task families.
08
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
2610.03710 · cs.RO, cs.AI · 2026-10-022610.03710 · cs.RO, cs.AI · 2026-10-02
EyeRobot 2.0 借鉴人类视觉,只用一台立体相机实现精细双臂操作:让两个「眼球」视角物理转动,把注视点对准场景中的 3D 注视目标,并在图像中心分配更多视觉 token(foveal 处理),把算力集中在任务相关特征上[40]。它要解决的是腕部相机的实际痛点:腕部相机虽然看得准,但增加硬件、线缆与易损点,且遮挡时失效。做法上把主动视觉注视与操作策略分层训练——先训练以目标物体为条件的低层注视伺服策略,再训练根据任务发出注视目标的目标选择器。对做机器人视觉的团队,值得借鉴的是用主动注视替代堆相机;限制是注视伺服的延迟与稳定性直接影响操作精度,动态场景下的鲁棒性仍需验证。
EyeRobot 2.0 borrows from human vision to enable fine-grained bimanual manipulation with only a single stereo camera: two eye viewpoints physically swivel to centre on a 3D fixation point, and the resulting images are processed foveally, allocating more visual tokens to the centre to focus compute on task-relevant features[40]. It addresses practical pain in wrist cameras, which see precisely but add hardware, cabling and failure points and fail under occlusion. Gaze and manipulation are trained hierarchically: a low-level gaze servoing policy conditioned on a goal object, then a target selector emitting fixation goals from the task. For robotics vision teams the transferable idea is replacing extra cameras with active gaze; the limit is that gaze latency and stability directly affect precision.
📅 覆盖口径
📅 Coverage
覆盖口径:北京时间 2026-10-04 00:00–23:00。
Coverage window: 2026-10-04 00:00–23:00 (UTC+8).
本文由自动化「AI资讯速递」工作流抓取公开信息后整理,评价与分析部分为个人观点,不构成投资或技术选型建议。
Compiled by an automated daily-trends workflow from public sources; the analysis reflects the author's personal views only.