状态: 活跃研究
受治理的 AI 系统
智能体治理、审批边界、人在回路控制、策略强制,以及可回滚、失败即关闭的执行。
范围
我们研究如何以单一受控执行边界约束自主与半自主 AI 系统:不可逆动作须经人工审批,工具访问基于策略授权。重点在于可审计、可回滚的机制,而非最大化自主性。
本方向的研究项目
- 状态: 活跃研究受治理的智能体系统
为 AI 智能体设置单一受控执行边界:不可逆动作须经人工审批,工具访问基于策略授权。
开放问题
- RQ-003设计札记哪一种治理原语应内置于智能体运行时之内,而非置于其之上?
发表
- A governed ReAct agent on a real local quantized model: benign tools run, injected side-effects are denied, memory is tenant-isolated
- Governed multi-sub-agent orchestration: plan, route, execute, verify, merge, audit, recover — with per-sub-agent isolation
- A scenario-aware routing policy with co-resident dense-small backends: the model bake-off as automatic runtime behavior
- A governed agent runtime: clarify the input, gate the outcome, compound the skill
- The sovereign product layer: integrating research engines into a coherent capability set
- Safe convergence for an experience-reuse product loop: belief provenance, a confirmation state machine, and adversarial design punch-throughs
- Risk classification cannot be static: from a Memory System to a Commitment System
- Commitment System: stands or falls — an adversarial integration of four punch-throughs
- The Commitment Decision Is Undecidable; the Commit Action Is Not
- Authority Is Not Liability: The Solvent-Backstop-Aware Runtime
- Capstone: the sovereign runtime is a fold of refused collapses terminating in stipulation by fiat
- Clarify Runtime: the model-free floor of the input gate
- Composing a model-free clarify-execute-verify loop into one sovereign orchestration module
- Skill Runtime: activation beats re-derivation, acquisition stays sovereign
- State Runtime: the agent's world is an event-sourced belief log
- Skill Factory: acquisition where the judge, not the teacher, is authority
- The epistemic guard: a hard gate against confident-wrong answers
- A sovereign ReAct agent harness under in-loop governance
- A sovereign decentralized bidding network (Ed25519 + a hash-chained ledger)
- 受治理 AI 智能体的单一执行边界:一份技术报告
