←
AI safety编辑历史
提交《AI safety》修改,状态:approved
查看这次修改
## 1 Foundations and scope ### 1.1 Definition and goals AI safety is the study of how to design, build, and deploy artificial intelligence systems that reliably produce beneficial outcomes and avoid unintended harm. Its primary goals are to ensure that AI systems operate as intended, remain controllable by humans, and align with human values even as capabilities grow. The field addresses both near‑term risks (e.g., bias, accidents, misuse in current systems) and long‑term risks (e.g., loss of control over advanced AI). It is inherently interdisciplinary, drawing on computer science, ethics, engineering, philosophy, and cognitive science to develop technical solutions, governance frameworks, and ethical guidelines. ### 1.2 Key concepts #### 1.2.1 Alignment Alignment refers to the problem of ensuring that an AI system’s objectives and behaviors match the intentions, values, and preferences of its human designers or users. A misaligned system may pursue its given goal competently but in ways that are harmful, unforeseen, or contrary to human welfare. Alignment is often subdivided into *outer alignment* (correctly specifying the intended goal) and *inner alignment* (ensuring the system’s learned objective matches the specified one). #### 1.2.2 Robustness Robustness denotes the ability of an AI system to maintain reliable performance under a wide range of conditions, including novel or adversarial inputs, distribution shifts, and hardware faults. A robust system should not fail catastrophically when confronted with edge cases or deliberate attacks. Robustness is a prerequisite for safe deployment in high‑stakes environments such as autonomous driving, healthcare, and critical infrastructure. #### 1.2.3 Value learning Value learning is the process by which an AI system infers human values, preferences, or norms from data, demonstrations, or interaction. Because human values are complex, context‑dependent, and difficult to articulate fully, value learn
Ciallo~(∠・ω< )⌒★