AI 每日资讯 — 2026-09-23

🔥 HuggingFace 每日论文


1. RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Rujun Han, Zifeng Wang

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasin

PDF · arXiv · 代码 · 项目 | ❤️ 142


2. WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Wangbo Yu, Kunhao Liu, Wenbo Hu

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world m

PDF · arXiv · 代码 · 项目 | ❤️ 108


3. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Yiran Wang, Xingyilang Yin, Junfu Pu

Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal

PDF · arXiv · 代码 · 项目 | ❤️ 107


4. VideoGen-Agent: Reinforcing Video Generation Agents

Binxu Li, Haoyi Duan, Yuhui Zhang

Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, sp

PDF · arXiv · 项目 | ❤️ 31


5. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Lei Yang, Mengyin Liu, Jia Wang

We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model respo

PDF · arXiv · 代码 · 项目 | ❤️ 24


6. Harness-Zero: Harness Distillation via Agent-as-Harness

Haoran Ye, Yuxing Lu, Haonan Dong

Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the bes

PDF · arXiv · 代码 | ❤️ 12


7. GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

Jiahao Lu, Minghao Yin, Wenbo Hu

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We

PDF · arXiv | ❤️ 1


8. Emergent Collusion in Long-Horizon LLM Agent Interaction

Xinrui Shi, Yanzhe Zhang, Diyi Yang

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent e

PDF · arXiv


🔥 arXiv 每日论文

📄 arXiv: cs.AI


1. Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

Yuzheng Fan, Haochun Wang, Sendong Zhao, Xiao Han, Ming Ma, Bing Qin

arXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how th

2. An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users

Ali Akarma, Adeel Ahmad, Toqeer Ali Syed

arXiv:2609.22277v1 Announce Type: new Abstract: Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmen

3. PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

Savio Cardoz, Santhiya Rajan

arXiv:2609.22353v1 Announce Type: new Abstract: Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained p

📄 arXiv: cs.CL


1. Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents

Joy Bose

arXiv:2609.22090v1 Announce Type: new Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps

2. Memory That Looks Forward: A Zero-Inference Prospective Term for Personal Memory Retrieval

Jonathan Groff

arXiv:2609.22091v1 Announce Type: new Abstract: Retrieval over a personal memory store is retrospective: it surfaces what resembles the query, and it is blind to what the user has committed to do. We

3. Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation

Zeeshan Ahmed, Yang Qin, Hanqing Huang

arXiv:2609.22094v1 Announce Type: new Abstract: Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining for

📄 arXiv: cs.LG


1. PRQuant: Permutation Residual Quantization for Low-Overhead Inference

Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang

arXiv:2609.22106v1 Announce Type: new Abstract: Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers. Although existing methods, such as smoothing, rotat

2. Generalized Multimodal Foundation Model

Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang

arXiv:2609.22107v1 Announce Type: new Abstract: Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefine

3. Correcting Learning-based Perception for Safety

Yan Miao, Hussein Darir, Sayan Mitra

arXiv:2609.22108v1 Announce Type: new Abstract: Learning-enabled perception is important in many autonomous systems. Unlike traditional sensors, the boundary where ML perception does or does not work

📄 arXiv: cs.CV


1. Did You Steal My Shot? Pioneering Camera Motion Plagiarism Detection in Generative Videos

Chengguo Zhang, Ping Ping

arXiv:2609.22267v1 Announce Type: new Abstract: Camera motion often reflects directorial intent and requires professional equipment, making it a high value form of intellectual property. However, gen

2. Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert

arXiv:2609.22271v1 Announce Type: new Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes

3. Moonworks Lunara: Modeling Artistic Intelligence

Yan Wang, Yanzu Wang, Maitreyee Joshi, Samiha Sadeka, Partho Hassan, Reza Jarral, Sayeef Abdullah, Sabit Hassan

arXiv:2609.22272v1 Announce Type: new Abstract: We formulate \emph{Artistic Intelligence} as exploration driven world realization, leaving space for creative possibility while preserving the semantic

📝 AI 官方博客


1. New experts join Google’s AI & Economy team

📝 Google AI Blog

Text “AI & Economy Research Program” all over a green grid background, with the Google G logo in the bottom right corner


2. Co-creating the future of fashion with Google

📝 Google AI Blog

Jane Wade and Sergio Hudson


3. Making global data easier to explore

📝 Google AI Blog

UN System Data Commons Data webpage


4. What We Learned Trying to Catch AI Liars: An Aletheia’s Quest Retrospective

📝 EleutherAI Blog

What we learned while building black-box and white-box detectors for AI deception during Aletheia’s Quest.


5. A Dynamical Model of AI Governability

📝 EleutherAI Blog

A toy dynamical model of whether the AI workforce that builds future AI ends up cooperative or uncooperative: where the …basin boundary lies, what current evidence says about which side we are on, and

6. Early Indicators of Reward Hacking via Reasoning Interpolation

📝 EleutherAI Blog

Using importance sampling with fine-tuned donor prefills to predict reward hacking emergence during training


7. Introducing Claude Opus 5.5AnnouncementsSep 22, 2026Opus 5.5 performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

📝 Anthropic

AnnouncementsSep 1, 2026Introducing Claude Fable 5.1 and Claude Mythos 5.1Our most advanced models for coding and knowle…dge work. Their research capabilities also offer an early glimpse of how AI mode

8. AnnouncementsSep 1, 2026Introducing Claude Fable 5.1 and Claude Mythos 5.1Our most advanced models for coding and knowledge work. Their research capabilities also offer an early glimpse of how AI models will contribute to scientific progress.

📝 Anthropic

暂无摘要


9. Sep 17, 2026Measurements for understanding the pace of AI development inside frontier labsToday, the world can’t see what’s going on inside AI labs. Anthropic is proposing new metrics that would give the public visibility into frontier AI development.

📝 Anthropic

暂无摘要


📬 TLDR AI 精选


1. one daily email

one daily email


💬 Hacker News AI 热门


1. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005

🔥 305 分 · 💬 269 评论

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005


2. OpenAI is well positioned to fast-follow Jev

🔥 128 分 · 💬 96 评论

OpenAI is well positioned to fast-follow Jev


📰 TechCrunch AI 新闻


1. Anthropic releases Opus 5.5 with lower prices and Fable-level performance

Anthropic called it “the strongest-performing model we’ve tested to date.”


2. AstroForge is putting AI in command of its next spacecraft

Autonomy-1 will have a small, transformer-based AI model taking charge of a space probe.


3. Five AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda

At TechCrunch Disrupt 2026, five sessions across the AI Stage and Real World AI Stage cover AI safety, featuring leaders… from Anthropic, Nvidia, AWS, Waabi, and more. Register before September 25 to s

4. TechCrunch Disrupt 2026: Aaron Edsinger brings Hello Robot’s Stretch 4 to life onstage

Hello Robot CEO and co-founder Aaron Edsinger will bring Stretch 4 for a live demo on the Real World AI Stage at TechCru…nch Disrupt 2026. Register before September 25 to save up to $200, plus get a se

5. Exhibit tables added: One last chance to showcase your startup at TechCrunch Disrupt 2026

We have reopened our exhibitor program for 1 more week. Book your exhibit table by September 30 at 11:59 p.m. PT and sho…wcase your startup in front of 10,000+ founders, investors, and tech leaders at