Changkun Ou
Science and art, life in between.科学与艺术,生活在其间。

Changkun Ou

Human-AI interaction researcher, engineer, and writer.人机交互研究者、工程师、写作者。

Bridging HCI, AI, and systems programming. Building intelligent human-in-the-loop optimization systems. Informed by psychology, sociology, cognitive science, and philosophy.连接人机交互、AI 与系统编程。构建智能的人在环优化系统。融合心理学、社会学、认知科学与哲学。

A Book on Bayesian Optimization一本关于贝叶斯优化的书

I wrote a book: Bayesian Optimization: From First Principles to Human Preferences. My PhD at LMU Munich was about human-in-the-loop systems, and one question kept coming back. Some things can only be judged by a person: whether a photo looks right, whether a 3D model still looks good after it is …

我写了一本书:《贝叶斯优化:从基本原理到人类偏好》。 我在慕尼黑大学读博时做的是人在回路(human-in-the-loop)的系统,有一个问题反复出现:有些东西只能靠人来判断,比如一张照片调得对不对,一个 3D 模型简化之后还好不好看,一个设计看着舒不舒服。这类问题没有公式可以拿来优化,只能去问人,而每问一次,都要花掉对方的时间和耐心。怎样用尽可能少的提问,找到一个好的设置?标准答案是贝叶斯优 …

Read More阅读更多 →

See You Down the Road后会有期

When I left Sixt, I sent a farewell letter to the people I worked with. This is that letter, with the names of people and internal systems left out. Those it thanks know who they are. Dear friends, colleagues, and everyone I have shared a project, a debate, a lunch, a late evening, a weekend, or an …

离开 Sixt 时,我给共事过的人写了一封告别信。下面就是这封信,略去了人名和内部系统的名字。信里感谢的每一位,自己会知道。 亲爱的朋友们、同事们,以及每一位和我一起做过项目、争论过问题、吃过午饭、熬过夜、过过周末、有过共同经历的人: 我写这封信,是为了给我在 Sixt 的这篇“论文”收尾。和每篇论文一样,最后留下来的也许不是正文的章节,而是致谢,所以我就从致谢写起。 我和 Sixt 的旅程开始于 …

Read More阅读更多 →

Trusting Trustworthiness: Zero Trust, Trust by Default, and the People Who Wrote the Software信任可信性:零信任、默认信任,以及写这个软件的人

“Perhaps it is more important to trust the people who wrote the software.” Ken Thompson, Reflections on Trusting Trust (1984) Suppose an actor’s values are genuinely positive. They do not seek domination, personal enrichment, or the suffering of others. They sincerely want to protect what is …

“或许更重要的是,信任写这个软件的人。” Ken Thompson,《Reflections on Trusting Trust》(1984) 假设某个行动者的价值观确实是好的:不求支配,不图私利,也不愿看到别人受苦;真心想守护有价值的东西,减少伤害,让世界比自己来时更好一些。那么,当某个具体的决定被证明是错的,这种内在的倾向,还能让人信任吗? 事故 2026 年 7 月,OpenAI 在内部做 …

Read More阅读更多 →
idea想法

Life Stage Axiom System人生阶段公理体系

Looking back over my life so far, I’ve arrived at three working “axioms” for this stage of it:

  • Axiom One (Keep an Exit): A position is worth what it pays plus the freedom to leave at any time, and the freedom counts for far more. Don’t take positions that lock you in.

  • Axiom Two (No Single Gatekeeper): Only enter systems where my feedback doesn’t depend on one particular person’s approval.

  • Axiom Three (Invest Where It’s Unclear): Invest in proportion to how unclear the goal is: the vaguer the goal, the more I put in.

The following content is generated by LLMs and may contain inaccuracies.

Axiom System for Current Life Stage

Context

This note attempts to distill a self-consistent set of decision axioms for “my current life stage”—a personalized system of action constraints. It spans personal decision theory, optionality, principal-agent problems, and the exploration-exploitation tradeoff. Its timeliness lies in compressing scattered life experience into three transferable rules: keeping an exit, no single gatekeeper, and investing where it’s unclear. The core tension: how to construct action principles in uncertain environments that both avoid structural lock-in and preserve growth opportunities. Notably, the author explicitly limits these to axioms “for the current stage,” suggesting a dynamic system that evolves with life circumstances rather than eternal truth.

Key Insights

  • Axiom One (Keep an Exit) is essentially the language of option value. The right to “leave anytime” corresponds financially to the early exercise value of an American option, whose weight far exceeds current actual returns—highly consistent with Nassim Taleb’s antifragility and optionality concepts (Taleb, Antifragile). The author’s “structural lock-in” maps to path dependence/lock-in in economics; David’s classic QWERTY keyboard analysis shows how suboptimal structures persist through lock-in (David, “Clio and the Economics of QWERTY”, American Economic Review 1985). Factoring exit costs into location valuation is a defensive reversal of the sunk cost fallacy.

  • Axiom Two (No Single Gatekeeper) avoids single-point veto power. “Feedback need not pass through one specific person’s confirmation” directly corresponds to gatekeeper risk and principal-agent problems in organizational theory (Jensen & Meckling, “Theory of the Firm”, Journal of Financial Economics 1976). When feedback loops depend on a single decision-maker, individuals become exposed to that person’s preferences, moods, and presence—aligned with the permissionless attribute sought by decentralized systems. By analogy: market prices need no one’s “approval” of your success, while bureaucratic promotion requires your superior’s signature. This axiom actually favors objective feedback systems (markets, code functionality, reader engagement) over subjective judgment systems (boss evaluations, committee scores).

  • Axiom Three (Invest Where It’s Unclear) inverts intuitive resource allocation. Common sense tilts toward “the clearer the goal, the more investment warranted”; the author proposes “investment intensity is proportional to goal ambiguity.” This echoes the exploration-exploitation tradeoff’s emphasis on exploration: in information-scarce, high-variance domains, marginal investment’s expected information value is higher (Sutton & Barto, Reinforcement Learning). It resonates with startup wisdom—the largest opportunities often exist in unclefined “blanks”; clear goals mean competition is saturated, excess returns arbitraged away. Understand this as a form of information arbitrage: ambiguity is mispricing, mispricing is opportunity. The counterargument: ambiguity may be mere noise rather than opportunity, requiring a discrimination mechanism between “unexploited vacancies” and “phantom illusions.”

  • Internal synergy of three axioms. These are not isolated: Axiom One reserves exit rights, Axiom Two ensures feedback isn’t monopolized, Axiom Three drives investment toward ambiguous terrain—together they form a low lock-in, high autonomy, exploration-biased action profile. This precisely describes the ideal ecological niche of independent creators, researchers, or early-stage founders, explaining why this axiom system appears in a researcher’s blog context.

Open Questions

  • Do Axiom Three (higher ambiguity → greater investment) and Axiom One (preserve anytime exit rights) create tension?—Overcommitting to highly ambiguous goals may itself accumulate hard-to-exit sunk costs and identity lock-in; how design a mechanism enabling deep exploration without becoming locked by exploration itself?

  • These axioms are explicitly limited to “current stage”—what signals trigger axiom revision? Is there a meta-axiom determining when to abandon “keeping an exit” for deliberate structural lock-in (when deep commitment’s compound returns exceed option value preservation)?

整理了一下自己的经历,给现阶段的自己总结出三条“公理”:

  • 公理一(留有退路):一个位置的价值,是它的实际回报加上随时可以离开的自由,而后者远比前者重要。不进入会把自己锁死的结构。
  • 公理二(不依赖单一评判者):只进入那些反馈不取决于某一个人认可的系统。
  • 公理三(向模糊处投入):投入的力度与目标的模糊程度成正比,目标越模糊,投入越多。

以下内容由 LLM 生成,可能包含不准确之处。

人生阶段公理体系

Context

这则笔记试图为"当前人生阶段的我"提炼一套自洽的决策公理——一种个人化的行动约束系统。它触及的领域横跨个人决策理论(personal decision theory)、期权思维(optionality)、激励与代理问题(principal-agent problems),以及探索-利用权衡(exploration-exploitation tradeoff)。之所以在当下值得记录,是因为它把散乱的人生经验压缩为三条可迁移的规则:留有退路、不依赖单一评判者、向模糊处投入。核心张力在于——如何在不确定环境中构造一套既能规避结构锁定(structural lock-in)、又能保留成长机会的行动准则。值得注意的是,作者明确将其限定为"目前阶段"的公理,暗示这是一个随人生阶段演化的动态体系,而非永恒真理。

Key Insights

  • 公理一(留有退路)本质上是期权价值的语言。 “随时能走"的权利在金融上对应美式期权(American option)的提前行权价值,其权重远大于当前实际收益,与 Nassim Taleb 提出的**反脆弱性(antifragility)**和 optionality 思想高度一致——保留选择权本身就是一种在不确定性中受益的结构(Taleb, Antifragile)。作者提到的"结构锁定"在经济学中即 path dependence / lock-in,David 关于 QWERTY 键盘的经典分析说明了次优结构如何因锁定而长期存续(David, “Clio and the Economics of QWERTY”, American Economic Review 1985)。将退出成本纳入位置价值的核算,是对沉没成本谬误的反向防御。

  • 公理二(不依赖单一评判者)是对单点否决权的规避。 “反馈不需要经过某个特定的人确认"直接对应组织理论中的gatekeeper 风险与代理问题(Jensen & Meckling, “Theory of the Firm”, Journal of Financial Economics 1976)。当反馈回路必须经由单一裁决者,个体就暴露于该裁决者的偏好、情绪与在场与否——这与去中心化系统追求的permissionless(无需许可)属性一脉相承。类比市场机制:公开市场的价格信号不需要任何人"批准"你的成功,而科层组织的晋升则依赖上级签字。这条公理实际上是在偏好客观反馈系统(市场、代码是否运行、读者是否阅读)而非主观裁决系统(老板评价、评委打分)。

  • 公理三(向模糊处投入)颠倒了直觉的投入分配。 常识倾向于"目标越清晰越值得投入”,而作者主张"投入强度正比于目标的模糊程度”。这与探索-利用权衡中对探索的偏重相呼应:在信息稀缺、结果方差大的模糊领域,边际投入的期望信息价值更高(Sutton & Barto, Reinforcement Learning)。它也与创业领域的观点共鸣——最大的机会往往存在于尚未被清晰定义的"空白"处,清晰的目标意味着竞争已充分、超额收益已被套利。可将其理解为一种信息套利:模糊即定价错误,定价错误即机会。潜在的反论是,模糊性也可能只是噪声而非机会,需要一个判别机制区分"未被开发的空缺"与"本就不存在的幻象"。

  • 三条公理的内在协同。 三者并非孤立:公理一保留退出权,公理二保证反馈不被单点垄断,公理三驱动向模糊地带投入——组合起来构成一个低锁定、高自主、偏探索的行动画像。这恰好描述了独立创作者、研究者或早期创业者的理想生态位,也解释了为何这套公理会出现在一位研究者的博客语境中。

Open Questions

  • 公理三(模糊度越高投入越大)与公理一(保留随时退出权)是否存在张力?——向高度模糊的目标重仓投入,本身可能积累难以退出的沉没成本与身份认同锁定,如何设计一个既深度探索又不被探索本身锁定的机制?

  • 这三条公理被明确限定为"目前阶段",那么触发公理更新的信号是什么?是否存在一条元公理,用来判定何时应当抛弃"留有退路"而主动选择结构锁定(例如深度承诺带来的复利回报超过了保留期权的价值)?

idea想法

Developing Taste Through Accumulated Experience通过累积经验培养品味

Tastes are accumulated from experience. Working on many different problems in the past teaches you what kinds of problems might be interesting in the future, or what kinds of things might be just barely possible by combining previous approaches. This can reveal open problems you might need to work on to achieve something magical or highly useful. Another way to gain experience is to write down a bunch of things you think might be important in the next 12 months. Maybe you pick one to work on, but then revisit and evaluate after 12 months—which of these other things actually proved important? Which ones did other people in the world create, and which ones haven’t been done yet? That can give you many more samples for developing your own taste-creation capability. That’s an important skill to have.

The following content is generated by LLMs and may contain inaccuracies.

Context

This idea sits at the intersection of research methodology, expertise development, and metacognition. It addresses a question rarely made explicit in scientific and creative training: how does one develop taste — the intuitive sense for which problems are worth pursuing and which combinations of ideas might yield something “magical or highly useful.” The claim is that taste is not innate but accumulated from experience working on a diversity of problems. That accumulated base teaches you (a) what future problems might be interesting, and (b) what might be just barely possible by cobbling together previous approaches plus a handful of open problems you’d still have to solve. The core practical insight is a technique for accelerating this accumulation: write down a list of things you think will be important in the next 12 months, work on one, and then return after 12 months to evaluate which predictions held — including which ideas other people in the world went out and created, and which no one has tackled yet. This retrospective scoring gives you “more samples” for your own taste-creation capability.

Key Insights

  • Taste as pattern recognition over accumulated cases. The framing that breadth of past problems teaches you what is “just barely possible” mirrors expertise research: expert intuition is compiled from a large library of encountered patterns, reliable only in domains with valid feedback structures. See Kahneman and Klein’s joint work on the conditions under which expert intuition can be trusted (Kahneman & Klein, “Conditions for Intuitive Expertise,” American Psychologist, 2009).

  • “Cobbling together previous approaches plus open problems." This describes research taste as a form of combinatorial search over an adjacent-possible frontier — recombining existing tools until something new becomes reachable. Kauffman’s notion of the “adjacent possible” captures why breadth expands what is barely-attainable (Stuart Kauffman, Investigations, 2000).

  • The 12-month prediction list as a calibration mechanism. Writing down predictions and revisiting them is a documented method for improving forecasting judgment; keeping records and scoring outcomes is central to Tetlock’s “superforecasting” findings, where deliberate feedback loops sharpen calibration far more than raw intelligence (Tetlock & Gardner, Superforecasting, 2015). The original idea generalizes this from probability calibration to taste calibration — evaluating not just “was I right” but “was this actually important / did the world need it.”

  • The three-way outcome scoring is the crucial refinement. The original proposes evaluating each written-down idea along distinct axes: (1) which ones actually seemed important in retrospect, (2) which ones others went out and created, and (3) which ones no one has done yet. This last category is the most valuable signal — an idea you flagged, still believe in, and that remains unbuilt is a candidate for a genuinely uncrowded opportunity. This resembles how investors and researchers hunt for correct-but-non-consensus bets.

  • Prediction as a substitute for slow lived experience. The explicit motivation is generating “more samples” faster than one problem-at-a-time work allows. This is a form of accelerated experiential learning: instead of waiting years to accumulate cases, you manufacture a batch of forecasts and let time adjudicate them cheaply. It parallels the practice of maintaining a decision journal to counter hindsight bias, since without a written record memory rewrites what you actually believed.

  • A structural caution. Betting on the near future favors ideas whose importance is visible within 12 months, potentially biasing taste toward incrementally legible problems and away from long-horizon bets whose payoff or even relevance won’t clarify for a decade. The technique may need multiple time-horizon lists to avoid short-termism.

Open Questions

  • If you keep 12-month prediction lists over many years, does your hit rate on “what was important” actually improve — or does the world’s fashion-driven attention (which ideas got hyped vs. genuinely mattered) contaminate the feedback signal and calibrate you toward trendiness rather than truth?

  • Can taste developed through recombination of your own past problems ever produce a paradigm shift, or does the “cobbling together” model structurally confine you to the adjacent possible — and what deliberate practice would push taste beyond it?

品味是从经验中累积而来的。在过去从事许多不同的问题能教会你什么样的问题在未来可能很有趣,或者什么样的东西可能刚好可以通过结合之前的方法来实现。这可以揭示你可能需要解决的一些开放问题,以实现令人惊艳的或非常有用的东西。获得经验的另一种方式是列出你认为在未来12个月内可能很重要的一堆事项。也许你选择其中一个来从事,但在12个月后重新审视并评估——这些其他事项中哪些实际上证明是重要的?世界上其他人创造了哪些,而哪些还没有被完成?这可以为你开发自己的品味创造能力提供更多样本。这是一项重要的技能。

以下内容由 LLM 生成,可能包含不准确之处。

背景

这个想法处于研究方法论、专业知识发展和元认知的交叉点。它针对一个在科学和创意训练中很少被明确提出的问题:人如何培养品味 — 一种直观的感觉,能够判断哪些问题值得追求,哪些想法的组合可能产生"魔法般的或高度有用的"结果。其主张是品味不是与生俱来的,而是从解决多样化问题的经验中积累而来。那些积累的基础教会你(a)未来哪些问题可能有趣,以及(b)通过拼凑之前的方法加上一些仍需解决的开放问题,什么是刚好可能的。核心实践洞察是加速这种积累的一种技术:列出你认为在接下来的12个月内会很重要的事物,专注于其中一个,然后在12个月后返回评估哪些预测成立 — 包括世界上其他人创造了哪些想法,以及哪些还没有人解决。这种回顾性评分为你的品味培养能力提供了"更多样本"。

关键洞察

  • 品味作为对积累案例的模式识别。 将过去问题的广度教会你什么是"刚好可能"的这一框架,与专业知识研究相呼应:专家直觉是从大量遇到的模式编译而来的,只有在具有有效反馈结构的领域中才可靠。见Kahneman和Klein的联合著作关于何时可以相信专家直觉的条件(Kahneman & Klein,《直觉专业知识的条件》,美国心理学家,2009)。

  • “拼凑之前的方法加上开放问题”。 这将研究品味描述为对邻近可能边界的组合搜索 — 重新组合现有工具,直到某些新东西变得可达。Kauffman的"邻近可能"观念捕捉了为什么广度扩展了什么是勉强可达成的(Stuart Kauffman,《调查》,2000)。

  • 12个月预测清单作为校准机制。 写下预测并重新审视它们是改进预测判断的一种被证实的方法;保留记录和评分结果是Tetlock的"超级预测"发现的核心,其中有意反馈循环比原始智力更能大幅提高校准(Tetlock & Gardner,《超级预测》,2015)。最初的想法将这从概率校准推广到品味校准 — 评估不仅"我是否正确",还有"这实际上是否重要/世界是否需要它"。

  • 三向成果评分是关键的改进。 最初的想法提议沿着不同的轴评估每个写下的想法:(1)哪些在回顾中实际上似乎很重要,(2)哪些其他人去创造了,以及(3)哪些还没有人做过。最后这个类别是最有价值的信号 — 你标记出来的、仍然相信的、且仍未实现的想法是真正不拥挤机会的候选。这类似于投资者和研究人员如何寻找正确但非共识的赌注。

  • 预测作为缓慢生活经验的替代。 明确的动机是生成"更多样本"的速度比逐个问题工作允许的要快。这是一种加速的经验学习形式:不是等待多年积累案例,而是制造一批预测并让时间廉价地来判决它们。它与维护决策日志的做法相似,以对抗事后聪慧偏见,因为没有书面记录,记忆会改写你实际相信的东西。

  • 一个结构性警告。 赌注近未来倾向于支持其重要性在12个月内可见的想法,可能使品味偏向于渐进式可理解的问题,而远离长期赌注,其收益甚至相关性在十年内都不会澄清。该技术可能需要多个时间视野清单来避免短期主义。

开放问题

  • 如果你在多年间保持12个月预测清单,你在"什么是重要的"上的命中率是否实际改善 — 或者世界以时尚为驱动的注意力(哪些想法被炒作vs.真正重要)是否污染反馈信号,使你校准到趋势而非真实?

  • 通过重新组合你自己的过去问题开发的品味能否产生范式转变,或者"拼凑"模型在结构上是否将你局限于邻近可能 — 什么刻意练习会将品味推向其之外?

idea想法

Illusion of General Technology and Evolution of Specialized Architecture通用技术幻觉与特化架构演进

This interview is very timely, and some of the insights within it are quite thought-provoking, particularly regarding the development of general-purpose technology. Many people are willing to believe that the “Bitter Lesson” claims general-purpose technology will ultimately triumph over specialized technology. However, in reality, whether we look at the historical experience of Moore’s Law or the current development of models, we discover that so-called general-purpose technologies are essentially illusions. Technologies that can truly scale are all moving toward specialization. For example, general-purpose computing has evolved into today’s heterogeneous computing, and Agent Harness design has evolved into a general design of full-stack models + inference engines + Agent Harness + Workload Scheduler. While high-level general abstractions have some value, a generalization that cannot grasp bottom-level details can never produce designs capable of bearing load and effectively encapsulating complexity.

https://www.youtube.com/watch?v=ffdR5fZTC5E

The following content is generated by LLMs and may contain inaccuracies.

Context

This note presents a counterintuitive argument against a YouTube interview, proposing that the claim “general-purpose techniques will ultimately defeat specialized ones” is largely an illusion. It directly challenges The Bitter Lesson, frequently cited in AI discourse—Rich Sutton’s assertion that methods relying on general computation and search/learning will eventually outperform those dependent on manual expert knowledge (Sutton, “The Bitter Lesson”). The note’s core tension is this: the path of scaling has actually been evolving toward specialization all along, with general abstraction effective only at high levels and incapable of truly encapsulating lower-layer complexity. This topic is particularly timely now, as the engineering stack for LLMs and Agents is rapidly differentiating from a “single large model” into a layered, heterogeneous system architecture.

Key Insights

  • The history of Moore’s Law is actually a history of specialization, not general-purpose victory. Single-core CPU performance gains stalled after Dennard scaling ended, forcing the industry toward heterogeneous computing—GPUs, TPUs, NPUs, DPUs, and more. Hennessy and Patterson explicitly identified Domain-Specific Architectures as the future in their Turing Award lecture, since general-purpose processors can no longer scale efficiently in energy terms (Hennessy & Patterson, “A New Golden Age for Computer Architecture”). This validates the note’s core observation: any technology that truly scales is moving toward specialization.

  • A critical clarification and rebuttal of the Bitter Lesson. Notably, Sutton’s “general-purpose” refers to general learning and search methods, not general hardware or system architecture. The note reveals a layer-mismatch commonly overlooked: even if algorithms pursue generality in learning, their physical and engineering implementation must be highly specialized to scale. In other words, generality resides in “what to optimize,” while specialization resides in “how to run it”—these are not contradictory. The “illusion” the note critiques is precisely this: mistaking method-level generality and incorrectly extrapolating it to the system level.

  • Layered specialization in Agent engineering stacks. The note traces evolution from early single Agent Harness designs to stratified generic design: full-stack model + inference engine + Agent Harness + Workload Scheduler. This aligns with actual infrastructure trends: at the inference engine layer, optimizations like vLLM’s PagedAttention target LLM-specific memory/throughput characteristics (Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention”); at the scheduling layer, Agent workloads (long-tail, multi-turn, heterogeneous tool invocations) demand specialized workload scheduling, not repurposed microservice schedulers. Each layer appears “generic” yet contains highly specialized internal implementations.

  • “High-level generic abstractions work, but cannot encapsulate lower-layer complexity”—the abstraction leak. This resonates deeply with Joel Spolsky’s Law of Leaky Abstractions (all non-trivial abstractions leak to some degree) (Spolsky, “The Law of Leaky Abstractions”). The note extends this software engineering principle to AI system architecture: a complexity encapsulation that can “bear load” must perceive and exploit lower-layer specialized details (memory hierarchy, operator fusion, hardware topology). Pure high-level generic interfaces cannot achieve this.

  • An implicit dialectical tension. The “full-stack + inference engine + Harness + Scheduler generic design” the note describes is itself a form of layered generality—generality hasn’t disappeared but retreated into “interface contracts for composing specialized modules.” This suggests the true answer may not be a binary “general vs. specialized” dichotomy, but rather a symbiotic structure where generality orchestrates and specialization executes, analogous to RISC’s simple generic instruction set carried by compiler/microarchitecture specialization.

Open Questions

  • If each layer’s “generic abstraction” must necessarily be backed by specialized implementation, where lies the sustainability boundary of system architecture evolution—could fragmentation of specialization eventually strangle scaling at some complexity threshold, cyclically recalling a new round of generic consolidation?

  • If the Bitter Lesson remains valid at the method level while the system level moves toward specialization, does this suggest an as-yet-unnamed “middle-layer law”—that true scaling viability depends neither on the most generic algorithm nor the most specialized hardware, but rather on orchestration-layer design that most effectively bridges abstraction leaks?

这个访谈很应景,里面有些认知很有感触。其中关于通用技术发展的讨论尤其值得关注。很多人愿意相信「Bitter Lesson」所说的通用技术最终会胜过非通用技术,但实际上,无论从摩尔定律的历史经验来看,还是从现在模型发展的角度来看,我们都会发现所谓的通用泛化技术本质上都只是一些错觉。真正能够规模化的技术都在朝着特化的方向推进。比如通用计算演化到如今的异构计算,比如 Agent Harness 的设计演化到全栈模型 + 推理引擎 + Agent Harness + Workload Scheduler 的通用设计。高层通用抽象有一定的作用,但没办法掌握底层细节的通用永远无法做出经得起压力、对复杂性进行有效封装的设计。

https://www.youtube.com/watch?v=ffdR5fZTC5E

以下内容由 LLM 生成,可能包含不准确之处。

Context

这条笔记针对一段访谈(YouTube 访谈)提出一个反直觉论点:所谓「通用泛化技术终将战胜专用技术」在很大程度上是一种幻觉。它直接挑战了 AI 领域被反复引用的 The Bitter Lesson——Rich Sutton 认为依赖通用计算与搜索/学习的方法长期会压倒依靠人工专家知识的方法(Sutton, “The Bitter Lesson”)。笔记的核心张力在于:规模化(scaling)的路径本身其实一直在朝「特化」演进,通用抽象只在高层有效,无法真正封装底层复杂性。这个议题此刻尤为应景,因为 LLM 与 Agent 的工程栈正快速从「单一大模型」分化为分层、异构的系统架构。

Key Insights

  • 摩尔定律的历史其实是一部「特化史」而非「通用胜利史」。 通用 CPU 的单核性能红利在 Dennard scaling 终结后停滞,产业被迫转向 GPU、TPU、NPU、DPU 等异构计算(heterogeneous computing)。Hennessy 与 Patterson 在图灵奖演讲中明确指出,未来属于领域特定架构(Domain-Specific Architectures),因为通用处理器已无法在能效上继续 scaling(Hennessy & Patterson, “A New Golden Age for Computer Architecture”)。这印证了笔记的核心观察:真正能规模化的技术都在朝特化推进。

  • 对 Bitter Lesson 的关键澄清与反驳。 值得注意的是,Sutton 论证的「通用」指的是通用的学习与搜索方法(method),而非通用的硬件或系统架构。笔记恰恰揭示了一个常被忽略的层次错位:即便算法层面追求通用学习,其物理与工程实现却必须高度特化才能 scale。换言之,通用性存在于「what to optimize」,特化性存在于「how to run it」——这两者并不矛盾,笔记所批判的「幻觉」正是把方法层的通用性错误外推到系统层。

  • Agent 工程栈的分层特化。 笔记提出从早期的 Agent Harness 单一设计,演化到全栈模型 + 推理引擎 + Agent Harness + Workload Scheduler 的分层通用设计。这与当前基础设施的实际趋势吻合:推理引擎层出现 vLLM 的 PagedAttention 等针对 LLM 内存/吞吐特性的专门优化(Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention”);调度层则需要针对 Agent 工作负载(长尾、多轮、工具调用异构)做专门的 workload scheduling,而非套用传统微服务调度。每一层看似「通用」,其内部实现却各自高度特化。

  • 「高层通用抽象有作用,但无法封装底层复杂性」——即抽象泄漏。 这一论断与 Joel Spolsky 的Law of Leaky Abstractions(所有非平凡的抽象在某种程度上都会泄漏)高度呼应(Spolsky, “The Law of Leaky Abstractions”)。笔记把这一软件工程原理推广到 AI 系统架构:一个真正管用的复杂性封装,必须能感知并利用底层的特化细节(内存层级、算子融合、硬件拓扑),纯粹的高层通用接口做不到这一点。

  • 一个隐含的辩证张力。 笔记描述的「全栈 + 推理引擎 + Harness + Scheduler 的通用设计」本身其实是一种分层通用性(layered generality)——通用性并未消失,而是退居为「组合特化模块的接口约定」。这暗示真正的答案可能不是「通用 vs 特化」的二元对立,而是通用性作为编排层、特化性作为执行层的共生结构,类似于 RISC 用简单通用指令集承载、由编译器/微架构做特化优化的思路。

Open Questions

  • 如果每一层的「通用抽象」内部都必然由特化实现支撑,那么系统架构的可持续演进边界在哪里——特化的碎片化(fragmentation)会不会在某个复杂度阈值反过来扼杀 scaling,从而周期性地召回一轮新的通用化整合?

  • Bitter Lesson 在方法层依然成立而系统层却走向特化,这是否意味着存在一条尚未被清晰命名的「中层定律」——即真正决定谁能 scale 的,不是最通用的算法也不是最专用的硬件,而是能最有效跨越抽象泄漏的编排层设计?

idea想法

Inexpressibility Traps in Formal Systems形式系统中的表达性陷阱

In mathematics, we can develop the intuition that two systems can prove exactly the same theorems yet differ in how quickly things get discovered in them. The asymmetry is not about power, but only about what is cheap to say. This is grounded in description complexity with a resource bound, since the whole point is that cost is bounded.

We switch frames when the gain beats the cost of switching. The trap is that even after the alternative frame has genuinely become better, if the switching cost dwarfs the gain, the expected time to switch grows exponentially. A community can know it’s in the worse frame and still never move.

The punchline is: if a frame makes the better frame unstateable, the switching cost isn’t high—it’s undefined. That is, you can’t price a move you can’t describe.

How do you find a better frame if a better frame is undefined?

The following content is generated by LLMs and may contain inaccuracies.

Context

This idea sits at the intersection of proof theory, algorithmic information theory, and the sociology of scientific paradigms. It distinguishes two properties of a formal system that are usually conflated: deductive power (which theorems are provable) and expressive economy (how cheaply a given truth can be stated and reached). Two systems can be extensionally identical — proving exactly the same theorems — yet differ radically in the discovery cost of any particular result. The claim is that this asymmetry “is not about power, only about what is cheap to say,” and that it can be made precise via description complexity with a resource bound, i.e. a Kolmogorov-style measure where cost is explicitly bounded rather than idealized to the incomputable limit. The framing then models frame-switching as a decision under cost: a community adopts an alternative representation only when the gain exceeds the switching cost. The pathology to explain: a community can know it occupies the inferior frame and still, in expectation, never move — and worse, a frame can render its superior alternative literally unstateable, making the switching cost not merely high but undefined.

Key Insights

  • Provable equivalence vs. speed-up is a real theorem, not a metaphor. Gödel’s speed-up phenomenon shows that adding an axiom (or moving to a stronger system) can shorten proofs of statements provable in both systems by a non-elementary amount — the shortest proof in the weaker system can be astronomically longer. This is the rigorous core of “same theorems, different discovery cost.” See Gödel’s 1936 note “Über die Länge von Beweisen” and its modern treatment in proof complexity (Buss, Handbook of Proof Theory). The asymmetry is genuinely about representation, not provability.

  • The resource-bounded framing is the right one. Plain Kolmogorov complexity is uncomputable, so an unbounded “cost of saying” would be undefined for the wrong reason. Time-bounded Kolmogorov complexity ($K^t$) and Levin’s notion of complexity ($Kt = K + \log t$) build the resource bound in directly, which matches the note’s insistence that “the whole point is that cost is bounded.” See Li & Vitányi, An Introduction to Kolmogorov Complexity and Levin’s Universal search (1973). This is the natural home for pricing “what is cheap to say” in a frame.

  • The exponential-lock-in dynamic has an economic analogue. The claim that “even after the alternative frame has genuinely become better, if the switching cost dwarfs the gain, the expected time to switch grows exponentially” mirrors path dependence and technological lock-in: QWERTY, VHS, and network-effect standards persist despite known superior alternatives. See W. Brian Arthur, “Competing Technologies, Increasing Returns, and Lock-In by Historical Events” (Economic Journal, 1989) and David’s QWERTY study. A community “knowing it’s in the worse frame and still never moving” is exactly a coordination equilibrium that is Pareto-dominated but individually stable.

  • This generalizes Kuhn without requiring incommensurability of truth. Kuhn’s paradigms differ in what problems are even seen; here the twist is sharper — the frames prove the same theorems, so there is no truth-level disagreement, only a cost-level one. This is closer to Kuhn’s “On the essential tension” between tradition and innovation than to full incommensurability. See Thomas Kuhn, The Structure of Scientific Revolutions (1962).

  • The punchline is the genuinely new move: undefined, not high, switching cost. If frame A cannot even express the object that names frame B’s advantage, then the gain term in the switch-decision is not a large number — it has no value in A’s vocabulary. “You can’t price a move you can’t describe.” This is a self-referential inexpressibility, adjacent to Tarski’s undefinability of truth (a language cannot express its own truth predicate) — Tarski, The Concept of Truth in Formalized Languages. The relevant deficit is not deductive incompleteness (Gödel) but expressive incompleteness: some concepts require an extension of the language before they can be reasoned about at all.

  • Search under undefined objectives is the operational crux. The closing question — “how do you find a better frame if a better frame is undefined?” — is the same structural problem faced in open-ended search and novelty-driven exploration, where the target cannot be specified in advance. Lehman & Stanley’s “Abandoning Objectives: Evolution Through the Search for Novelty Alone” (Evolutionary Computation, 2011) argues that when the objective is deceptive or unstateable, objective-driven search fails and only novelty/diversity pressure discovers stepping stones. This suggests the escape route is not optimization within the frame but exploratory expansion of the frame’s vocabulary.

Open Questions

  • If a frame’s advantage is unstateable from inside the current frame, is the only reliable discovery mechanism a blind expansion of expressive vocabulary (analogous to novelty search or conservative axiom extension) — and can we characterize which vocabulary extensions are “safe” enough that a community would tolerate the switching cost before the payoff is visible?

  • Can the transition from “high but finite” to “undefined” switching cost be formalized as a phase boundary — e.g., in terms of whether the superior frame is a conservative extension (payoff pricable) versus a non-conservative reframing (payoff requires the new language to even be named)?

在数学中,我们可以形成这样的直觉:两个系统可能能够证明完全相同的定理,但在其中发现事物的速度上有所不同。这种不对称性不在于能力,而仅仅在于什么是便于表达的。这是基于描述复杂性和资源约束的,因为重点正是成本被限制了。

当收益超过切换成本时,我们会切换框架。陷阱在于,即使替代框架确实已经变得更好,如果切换成本远超收益,预期的切换时间会指数增长。一个社区可能知道自己处于更差的框架中,却永远无法移动。

关键之处在于:如果一个框架使更好的框架无法表述,那么切换成本就不是高,而是未定义的。也就是说,你无法为一个无法描述的举动定价。

如果更好的框架是未定义的,你如何找到一个更好的框架呢?

以下内容由 LLM 生成,可能包含不准确之处。

上下文

这个想法位于证明论、算法信息论和科学范式社会学的交叉点。它区分了一个形式系统的两个通常被混为一谈的性质:演绎能力(哪些定理可证)和表达经济性(陈述和得出给定真理的成本有多低)。两个系统可以外延等价——证明完全相同的定理——却在任何特定结果的发现成本上有根本差异。其主张是这种不对称"不是关乎能力,仅仅是关乎什么说起来很便宜",并且可以通过资源受限的描述复杂性精确化,即一种Kolmogorov风格的度量,其中成本被明确地限制而不是理想化为不可计算的极限。该框架随后将框架切换建模为成本下的决策:当收益超过切换成本时,一个共同体才采用替代表示。要解释的病理现象:一个共同体可能知道它处于劣势框架,但在期望意义上永远不会移动——更糟的是,一个框架可能使其优越的替代方案字面上无法表述,使切换成本不仅高,而是未定义的。

关键洞察

  • 可证等价性与加速是实实在在的定理,而非比喻。 哥德尔的加速现象表明,添加公理(或转向更强的系统)可以将在两个系统中都可证的陈述的证明缩短非初等级别的量——较弱系统中最短的证明可能是天文数字般长的。这是"相同定理、不同发现成本"的严格核心。参见哥德尔1936年的论文《关于证明的长度》和其在证明复杂性中的现代处理(Buss,《证明论手册》)。这种不对称性确实关乎表示,而非可证性。

  • 资源受限框架是正确的。 平白的Kolmogorov复杂性是不可计算的,所以未受限的"说的成本"会因错误的原因而未定义。时间受限Kolmogorov复杂性($K^t$)和Levin的复杂性概念($Kt = K + \log t$)直接将资源受限内置其中,这与该论述坚持"整个要点是成本是受限的"相符。参见Li & Vitányi,《Kolmogorov复杂性导论》和Levin的《通用搜索》(1973)。这是为框架中"说起来便宜的事"定价的自然位置。

  • 指数锁定动态有一个经济学类似物。 “即使替代框架确实变得更好,若切换成本远超收益,切换的预期时间呈指数增长"的主张反映了路径依赖和技术锁定:QWERTY、VHS和网络效应标准尽管有已知的更优替代方案却依然存在。参见W. Brian Arthur,《竞争技术、递增收益和历史事件锁定》(《经济学杂志》,1989)和David关于QWERTY的研究。一个共同体"知道自己处于更差框架而仍然永不移动"正是一个Pareto次优但个体上稳定的协调均衡。

  • 这概括了Kuhn而无需真理的不可公度性。 Kuhn的范式在看到的问题上有所不同;这里的转折更尖锐——框架证明相同的定理,所以没有真理层面的分歧,仅有成本层面的。这更接近Kuhn的《论本质张力》(传统与创新之间)而非完全的不可公度性。参见Thomas Kuhn,《科学革命的结构》(1962)。

  • 决定性的新颖之处是:未定义而非高切换成本。 如果框架A甚至无法表述命名框架B优势的对象,那么切换决策中的收益项不是一个大数字——它在A的词汇中没有价值。“你无法为一个你无法描述的举动定价。“这是一种自指的不可表述性,与Tarski的真理不可定义性相邻(一种语言无法表述自身的真理谓词)——Tarski,《形式化语言中真理的概念》。相关的缺陷不是演绎不完全性(哥德尔),而是表达不完全性:某些概念在语言被扩展之前根本无法被推理。

  • 未定义目标下的搜索是操作上的关键。 结尾问题——“如果一个更好的框架是未定义的,你如何找到它”——与开放式搜索和新颖性驱动探索面临的结构性问题相同,其中目标无法提前指定。Lehman & Stanley的《放弃目标:通过独自寻求新颖性的进化》(《进化计算》,2011)论证当目标具有欺骗性或无法陈述时,目标驱动搜索失败,仅新颖性/多样性压力才能发现踏脚石。这表明逃脱路线不是框架内的优化,而是框架词汇的探索性扩展。

开放问题

  • 如果框架的优势从当前框架内无法陈述,是否唯一可靠的发现机制是盲目扩展表达词汇(类似于新颖性搜索或保守公理扩展)——我们能否刻画哪些词汇扩展是"足够安全的”,使得一个共同体愿意承受切换成本,即便在收益可见之前?

  • “高但有限"到"未定义"切换成本的转变能否被形式化为一个相位边界——例如,根据优越框架是否是保守扩展(收益可定价)与非保守重构(收益需要新语言才能被命名)?

idea想法

Human Generalization Over Token Accumulation人类泛化能力 vs 代币积累

We should never forget that the strongest aspect of human intelligence is our generalization and sample efficiency. Some people value and invest years of practice or large amounts of token consumption these days as a form of endorsement. That’s fair and does provide some degree of safety and establishes a baseline; but to generalize and grow exponentially, all you need is good intuition and curiosity. Most of this comes from pre-training—that is, early-stage education and environmental opportunities.

The following content is generated by LLMs and may contain inaccuracies.

Context

This idea sits at the intersection of cognitive science, machine learning theory, and the epistemology of expertise. It contrasts two models of intelligence: the human strength of generalization sample efficiency — learning powerful abstractions from very few examples — versus the increasingly dominant industry metric of token accumulation (years of practice, or literally the number of training tokens “burned”). The tension it addresses is timely: as large language models scale by consuming trillions of tokens, there’s a cultural drift toward valuing sheer accumulation as a proxy for competence and endorsement. The note argues this accumulation gives safety and a bar, but that exponential growth comes instead from good intuition and curiosity, which are largely shaped in a “pre-training” phase analogous to early-stage education and environment.

Key Insights

  • Sample efficiency is humanity’s signature advantage. Humans (and children especially) generalize from a handful of examples, where deep learning systems often require orders of magnitude more data. This gap is a central theme in Lake, Ullman, Tenenbaum & Gershman, “Building Machines That Learn and Think Like People”, who argue human learning leverages compositionality, causal models, and learning-to-learn rather than brute pattern accumulation.

  • The “pre-training” analogy is apt but double-edged. In ML, pre-training on broad data builds priors that make downstream few-shot learning efficient — the argument in the note that intuition/curiosity “comes from pre-training, aka early stage education and environment.” This mirrors developmental findings that early environment shapes later learning capacity (see the “learning to learn” or meta-learning framing in Thrun & Pratt, Learning to Learn). The double edge: if early priors are impoverished, the generalization advantage never fully develops — echoing environmental effects on cognitive development.

  • Curiosity as an intrinsic driver of efficient learning. The claim that curiosity fuels exponential growth is supported by work on intrinsic motivation and curiosity-driven exploration, e.g. Pathak et al., “Curiosity-driven Exploration by Self-supervised Prediction”, where prediction-error-based curiosity dramatically improves learning without external reward. Curiosity effectively selects high-information samples, boosting effective sample efficiency.

  • Accumulation buys a floor, not a ceiling. The note’s concession is important: years of practice / tokens burned provide “some degree of safety and a bar.” This aligns with the deliberate-practice literature (Ericsson et al.), which shows accumulation reliably produces competence — but competence and generalizing breakthroughs are not the same axis. Endorsement systems (credentials, seniority, benchmark token counts) reward the reliable floor because it is measurable, not because it captures the intuition that produces leaps.

  • The scaling-law counterpoint. The industry’s bet on token accumulation is not irrational: Kaplan et al., “Scaling Laws for Neural Language Models” and the compute-optimal reweighting in Hoffmann et al., “Chinchilla” show predictable capability gains from more tokens/compute. The note’s implicit critique is that these are linear-to-power-law returns on a system lacking the architectural priors that make human learning exponentially efficient per example — capability rises, but per-token efficiency does not approach human levels.

Open Questions

  • If human-level generalization stems from priors laid down in a “pre-training” developmental window, can we deliberately engineer curiosity and intuition priors into learning systems — or into education — rather than relying on token/practice accumulation to eventually approximate them?
  • Endorsement mechanisms reward accumulation because intuition is hard to measure; what would a credible signal for sample-efficient generalization look like, and would trusting it change who and what we value in both AI and human expertise?

我们永远不应该忘记,人类智能最强大的地方在于我们的泛化能力和样本效率。如今,有些人重视并投入多年的实践或大量的代币消耗作为一种认可形式。这是公平的,确实提供了某种程度的安全性并建立了基准;但要实现泛化并指数级增长,你只需要良好的直觉和好奇心。这些大多来自预训练——也就是说,早期教育和环境机会。

以下内容由 LLM 生成,可能包含不准确之处。

背景

这一想法处于认知科学、机器学习理论和专业知识认识论的交叉点。它对比了两种智能模型:人类的泛化样本效率优势——从极少数例子中学习强大的抽象——与日益主导的行业衡量标准令牌积累(多年的实践,或字面上"消耗"的训练令牌数量)。它涉及的张力具有时代意义:随着大型语言模型通过消耗数万亿令牌进行扩展,存在一种文化漂移,将单纯的积累作为能力和认可的代理。该笔记论证这种积累提供安全性和基准,但指数增长实际上来自良好的直觉和好奇心,这些主要在"预训练"阶段形成,类似于早期教育和环境。

核心见解

  • 样本效率是人类的标志性优势。 人类(特别是儿童)从少数几个例子进行泛化,而深度学习系统通常需要数量级更多的数据。这个差距是Lake、Ullman、Tenenbaum & Gershman 的《构建像人一样学习和思考的机器》的中心主题,他们论证人类学习利用组合性、因果模型和学会学习,而不是蛮力模式积累。

  • “预训练"类比是恰当的,但有双重性。 在机器学习中,在广泛数据上的预训练建立了先验,使下游少量样本学习变得高效——笔记中论证直觉/好奇心"来自预训练,即早期教育和环境”。这反映了发展研究的发现,即早期环境塑造后来的学习能力(参见Thrun & Pratt 的《学会学习》中的"学会学习"或元学习框架)。双重性在于:如果早期先验不足,泛化优势永远无法完全发展——这呼应了环境对认知发展的影响。

  • 好奇心作为高效学习的内在驱动力。 好奇心推动指数增长的主张得到了内在动机和好奇心驱动探索研究的支持,例如Pathak 等人的《通过自监督预测进行好奇心驱动的探索》,其中基于预测误差的好奇心在没有外部奖励的情况下大幅改进学习。好奇心有效地选择高信息样本,提高有效样本效率。

  • 积累购买底线,而非天花板。 笔记的让步很重要:多年实践/消耗的令牌提供"某种程度的安全性和基准"。这与刻意练习文献(Ericsson 等人)相一致,其显示积累可靠地产生能力——但能力和推广突破不是同一个维度。认可系统(证书、资历、基准令牌计数)奖励可靠的底线是因为它是可测量的,而不是因为它捕捉了产生飞跃的直觉。

  • 缩放定律的反驳。 行业对令牌积累的押注并非不理性的:Kaplan 等人的《神经语言模型的缩放定律》和Hoffmann 等人的《Chinchilla》中的计算最优再加权显示了从更多令牌/计算获得的可预测能力增益。笔记的隐含批评是这些是线性到幂律的回报,作用在缺乏使人类学习在每个例子上指数级高效的架构先验的基层——能力上升,但每令牌效率并未接近人类水平。

悬而未决的问题

  • 如果人类水平的泛化源于在"预训练"发展窗口中建立的先验,我们是否可以故意将好奇心和直觉先验工程化到学习系统中——或融入教育——而不是依赖令牌/实践积累最终近似它们?
  • 认可机制奖励积累是因为直觉难以测量;样本高效泛化的可信信号会是什么样子,相信它会改变我们在人工智能和人类专业知识中重视谁和什么吗?
idea想法

Balancing Agent Autonomy and Verification Cost平衡智能体自主性与验证成本

Agent economics is fundamentally about finding the right “autonomy affordance” in an organization.

When we allocate more compute, we can accelerate agent autonomy for execution. But the cost of human verification also rises sharply. The challenge is not only checking whether the output is correct, but actually understanding what happened, why it happened, and whether we can responsibly trust it, even when things are happening fully autonomously. On the other hand, too little compute without sufficient autonomy brings us back to excessive manual effort, where humans remain the main bottleneck.

So the actual game in this agent economic model is to engineer the correct level of autonomy where an organization of people can understand and verify outcomes within an allocated compute budget, while still meeting demand from the market they serve. This relies on continuous improvements in model capability, engineered economic loops, and the people who can operate based on this model.

The following content is generated by LLMs and may contain inaccuracies.

Context — This idea sits at the intersection of AI agent design, organizational economics, and human-AI trust calibration. As inference-time compute becomes cheaper and agentic systems (autonomous LLM-driven workflows) become capable of long-horizon execution, a new bottleneck emerges: not the production of work, but its verification. The note frames this as finding the right “autonomy affordance” — the level of delegated autonomy an organization can absorb given its capacity to understand, verify, and responsibly trust agent outputs within a fixed compute budget. The tension is bidirectional: too much autonomy inflates human verification cost sharply (checking correctness, but also comprehending what happened, why, and whether it can be trusted); too little autonomy reverts to manual effort where humans are the bottleneck.

Key Insights

  • Verification cost, not generation cost, becomes the binding constraint. This echoes the long-standing intuition in complexity theory that verifying a solution can be easier than producing it (the P vs NP asymmetry), but the note flips the practical concern: when generation is cheap and autonomous, human verification becomes expensive because it requires reconstructing context the agent traversed. This is related to the “oversight tax” discussed in work on scalable oversight (Amodei et al., Concrete Problems in AI Safety).

  • Understanding ≠ checking correctness. The note makes a sharp distinction between verifying an output is correct and understanding what/why happened well enough to responsibly trust it. This maps onto the interpretability and process-vs-outcome supervision debate — e.g. rewarding correct reasoning traces rather than just correct answers (Lightman et al., Let’s Verify Step by Step, OpenAI). Trust requires legibility of process, not just accuracy of result.

  • The “autonomy affordance” as an economic equilibrium. The note reframes agent deployment as an optimization: engineer the autonomy level where an organization of people can understand and verify outcomes within an allocated compute budget while still meeting market demand. This is a three-variable balancing act — model capability, an engineered economic loop, and human operators trained to work within it. This resonates with the concept of “human-AI complementarity” and comparative advantage in task allocation (Dell’Acqua et al., Navigating the Jagged Technological Frontier, HBS).

  • Compute allocation as a governance lever. Throwing more compute accelerates execution autonomy but does not automatically fund the verification side of the ledger. The implicit claim is that verification capacity must scale alongside — otherwise organizations accumulate un-auditable autonomous output. This connects to scalable oversight proposals like debate and recursive reward modeling (Irving et al., AI Safety via Debate; Leike et al., Scalable agent alignment via reward modeling).

  • The two failure modes are symmetric. Under-autonomy keeps humans as the throughput bottleneck (defeating the point of automation); over-autonomy shifts the bottleneck to verification and trust (defeating accountability). The “game” is locating the point between them — and critically, that point moves as model capability improves, so it is a dynamic equilibrium, not a fixed setting.

  • Organizational readiness as a co-requisite. The note insists the model requires “people who can operate based on this model” — implying that the human operators' skill in interpreting, spot-checking, and trusting agent output is itself part of the affordance. Autonomy is not purely a property of the agent but a joint property of agent + organization.

Open Questions

  • Can verification cost be amortized — e.g. by having agents produce structured, auditable rationales or by delegating verification to other (cheaper, adversarial) agents — so that the autonomy affordance expands without proportionally growing human oversight burden? At what point does agent-verifying-agent become circular rather than trust-building?

  • If the optimal autonomy level shifts every time model capability jumps, how should an organization design its verification processes and operator skills to be robust to that drift rather than continuously re-engineered — and who bears the transition cost when the equilibrium moves?

智能体经济学的核心在于在组织中找到正确的"自主性承载量"。

当我们分配更多算力时,可以加快智能体执行的自主性。但人工验证的成本也会急剧上升。挑战不仅在于检查输出是否正确,还要实际理解发生了什么、为什么发生,以及我们是否能够负责任地信任它,即使所有事情都在完全自主的情况下发生。另一方面,算力不足而自主性不足会导致过度的手工操作,使人类仍然成为主要瓶颈。

因此,在这个智能体经济模型中,实际的博弈是要找到正确的自主性水平,使得组织内的人员能够在分配的算力预算范围内理解和验证结果,同时仍然满足所服务市场的需求。这取决于模型能力的持续改进、经过设计的经济循环,以及能够基于这一模型开展运营的人员。

以下内容由 LLM 生成,可能包含不准确之处。

背景 — 这个想法位于AI代理设计、组织经济学和人-AI信任校准的交叉点。随着推理时计算成本下降,代理系统(自主型LLM驱动的工作流)具备长期执行能力,一个新的瓶颈出现了:不是工作的生产,而是其验证。该笔记将其框架化为找到正确的**“自主权承载能力”**——在组织有能力理解、验证和负责任地信任代理输出的固定计算预算范围内,可以委派的自主权水平。紧张关系是双向的:过多的自主权会急剧增加人工验证成本(检查正确性,还要理解发生了什么、为什么,以及是否值得信任);太少的自主权则回到手工劳动,人类成为瓶颈。

关键见解

  • 验证成本而非生成成本成为约束因素。 这呼应了复杂性理论中的长期直觉——验证一个解可能比生成它更容易(P与NP不对称),但该笔记翻转了实际关注点:当生成廉价且自主时,人工验证变得昂贵,因为它需要重建代理所经历的上下文。这与在可扩展监督工作中讨论的"监督税"相关(Amodei等人,AI安全的具体问题)。

  • 理解≠检查正确性。 该笔记在验证输出是否正确和充分理解发生了什么/为什么以负责任地信任它之间做出了尖锐区分。这映射到可解释性和过程对比结果监督的争论——例如,奖励正确的推理轨迹而非仅奖励正确答案(Lightman等人,让我们逐步验证,OpenAI)。信任需要过程的可读性,而非仅仅结果的准确性。

  • “自主权承载能力"作为经济均衡。 该笔记将代理部署重新框架化为一个优化问题:工程化自主权水平,使得一个人员组织可以在分配的计算预算内理解和验证结果,同时仍满足市场需求。这是一个三变量平衡行为——模型能力、工程化的经济循环,以及训练有素在其中运作的人类操作员。这与"人-AI互补性"概念和任务分配中的比较优势相呼应(Dell’Acqua等人,在参差不齐的技术前沿上航行,哈佛商学院)。

  • 计算分配作为治理杠杆。 增加更多计算加速执行自主权,但不会自动为验证端的分类账提供资金。隐含的说法是,验证能力必须随之扩展——否则组织会积累不可审计的自主输出。这与可扩展监督提案相连接,如辩论和递归奖励建模(Irving等人,通过辩论进行AI安全;Leike等人,通过奖励建模进行可扩展的代理对齐)。

  • 两种失败模式是对称的。 自主权不足使人类成为吞吐量瓶颈(违背自动化的目的);自主权过度将瓶颈转移到验证和信任(违背问责制)。“游戏"是定位它们之间的点——更关键的是,该点随着模型能力改进而移动,所以它是一个动态均衡,而非固定设置。

  • 组织准备就绪作为共同前提。 该笔记坚持模型需要"能够基于此模型运作的人”——隐示人类操作员在解释、抽样检查和信任代理输出方面的技能本身是承载能力的一部分。自主权不纯粹是代理的属性,而是代理+组织的联合属性。

开放性问题

  • 验证成本能否被摊销——例如,通过让代理生成结构化、可审计的理由,或者通过委派验证给其他(更廉价、对抗性的)代理——使得自主权承载能力扩展而无需成比例地增加人工监督负担?在什么时点代理验证代理会变成循环而非信任建立?

  • 如果最优自主权水平在每次模型能力跃升时都改变,组织应如何设计其验证流程和操作员技能以对这种漂移保持稳健而非持续重新设计——以及当均衡移动时谁承担过渡成本?

Why High-Output Systems Are Often the First to Stop Growing为什么最高产的系统往往最先停止成长

On two kinds of novelty: the kind that compounds and the kind that only piles up. “The limits of my language mean the limits of my world.” – Wittgenstein, Tractatus 5.6 For a while it looked like progress. For a week, an AI agent pipeline I ran kept shipping. A commit landed roughly every hour: …

两种“新”:一种会复利,一种只会越堆越多。 我语言的边界,就是我世界的边界。 – 维特根斯坦《逻辑哲学论》5.6 有一段时间,它看上去一直在进步。 我跑的一条 AI Agent 流水线,连续一周都在交付:差不多每小时一个 commit,修修小问题,做点小改进,提交记录看上去生机勃勃。可产品本身一点也没有变大1。它没出故障,也没闲着,只是一直在用自己已经会的那些东西打转。 这种错觉并不少 …

Read More阅读更多 →