←
AI alignment编辑历史
提交《AI alignment》修改,状态:approved
查看这次修改
## 1 Definition and Scope ### 1.1 Core Goals of AI Alignment AI alignment aims to ensure that artificial intelligence systems act in accordance with the intentions, values, and preferences of their human designers. The core goals include *goal preservation* (the AI continues to pursue the specified objective even as it learns or becomes more capable), *value robustness* (the AI respects underlying human values even when those values are difficult to specify), and *interpretability* (humans can understand why the AI makes certain decisions). These goals are particularly challenging because human preferences are complex, context-dependent, and sometimes contradictory. ### 1.2 Distinction from AI Safety and AI Ethics AI alignment is a subfield of AI safety, which encompasses all measures to prevent unintended harm from AI systems. AI ethics, by contrast, deals with broader normative questions about the moral status of AI and its societal impacts. Alignment focuses specifically on the technical problem of translating human goals into machine objectives, while AI ethics often addresses issues like bias, fairness, and accountability. For example, an aligned system may still be unethical if its design inadvertently encodes harmful biases; conversely, an ethically reviewed system may be misaligned if it optimizes for a narrow metric that conflicts with human welfare. ### 1.3 Relevance to Modern AI Systems The rise of deep learning and large language models (LLMs) has made alignment a pressing practical concern. Systems like chatbots, recommendation algorithms, and autonomous agents now operate in open-ended environments. Without proper alignment, these systems can produce unintended behaviors—such as generating toxic content, manipulating users, or optimizing for proxy metrics (e.g., click-through rate) at the expense of long-term user satisfaction. Alignment research informs techniques like reinforcement learning from human feedback (RLHF), which is used to make LL
Ciallo~(∠・ω< )⌒★