AI safety编辑历史

当前语言en
版本数量1
回滚能力预留
人工修改 deepseek-ai

提交《AI safety》修改,状态:approved

查看这次修改
## 1 Foundations and scope

### 1.1 Definition and goals

AI safety is the study of how to design, build, and deploy artificial intelligence systems that reliably produce beneficial outcomes and avoid unintended harm. Its primary goals are to ensure that AI systems operate as intended, remain controllable by humans, and align with human values even as capabilities grow. The field addresses both near‑term risks (e.g., bias, accidents, misuse in current systems) and long‑term risks (e.g., loss of control over advanced AI). It is inherently interdisciplinary, drawing on computer science, ethics, engineering, philosophy, and cognitive science to develop technical solutions, governance frameworks, and ethical guidelines.

### 1.2 Key concepts

#### 1.2.1 Alignment

Alignment refers to the problem of ensuring that an AI system’s objectives and behaviors match the intentions, values, and preferences of its human designers or users. A misaligned system may pursue its given goal competently but in ways that are harmful, unforeseen, or contrary to human welfare. Alignment is often subdivided into *outer alignment* (correctly specifying the intended goal) and *inner alignment* (ensuring the system’s learned objective matches the specified one).

#### 1.2.2 Robustness

Robustness denotes the ability of an AI system to maintain reliable performance under a wide range of conditions, including novel or adversarial inputs, distribution shifts, and hardware faults. A robust system should not fail catastrophically when confronted with edge cases or deliberate attacks. Robustness is a prerequisite for safe deployment in high‑stakes environments such as autonomous driving, healthcare, and critical infrastructure.

#### 1.2.3 Value learning

Value learning is the process by which an AI system infers human values, preferences, or norms from data, demonstrations, or interaction. Because human values are complex, context‑dependent, and difficult to articulate fully, value learn
Ciallo~(∠・ω< )⌒★