| 注册
首页|期刊导航|心理学报|大语言模型的人格化对齐及其对道德判断的影响

大语言模型的人格化对齐及其对道德判断的影响

李昌锦 焦丽颖 陈圳 许恒彬 吴胜涛 许燕

心理学报2026,Vol.58Issue(7):1237-1253,中插1-中插3,20.
心理学报2026,Vol.58Issue(7):1237-1253,中插1-中插3,20.DOI:10.3724/SP.J.1041.2026.1237

大语言模型的人格化对齐及其对道德判断的影响

Personality-based alignment of large language models and its impact on moral judgment

李昌锦 1焦丽颖 2陈圳 1许恒彬 1吴胜涛 3许燕1

作者信息

  • 1. 北京师范大学心理学部||应用实验心理北京市重点实验室||心理学国家级实验教学示范中心[北京师范大学],北京 100875
  • 2. 北京林业大学人文社会科学学院心理学系,北京 100083
  • 3. 吉林大学哲学社会学院哲学系,长春 130012
  • 折叠

摘要

Abstract

With the advent of the human-machine symbiosis era,the ethical dilemmas and algorithmic biases of large language models(LLMs)have triggered widespread societal concerns.Guiding artificial intelligence(AI)toward beneficial development has thus become an urgent and challenging imperative.This research explores the impact of personality-based alignment grounded in the HEXACO personality model on the moral judgment of LLMs.Specifically,the study aims to verify whether LLMs can effectively achieve personality-based alignment through prompting and to systematically evaluate how such alignment influences utilitarian tendencies in LLMs compared to humans across various moral dilemmas.By leveraging established psychological frameworks,this research seeks to provide a scientific basis for constructing controllable and ethical AI alignment strategies. Study 1 tested GPT-3.5,GPT-4,and ERNIE 3.5 using personality prompts targeting each HEXACO domain separately at high,low,and baseline levels,integrated with different gender roles.Manipulation checks were conducted using two distinct methods:a quantitative personality assessment using the HEXACO-60 scale and a qualitative personality story-writing task rated by independent human evaluators.Study 2 utilized a set of standardized moral dilemmas to assess utilitarian versus deontological choices in both LLMs and human participants.Human data were categorized into high and low personality groups for comparison,while the LLMs performed the same moral judgment tasks under various personality settings to identify shifts in decision-making patterns. The results of Study 1 confirmed the feasibility of personality-based alignment,demonstrating that LLMs can dynamically represent HEXACO personality traits through prompts.Among the LLMs tested,GPT-4 exhibited superior instruction-following capabilities and more distinct trait differentiation than GPT-3.5 and ERNIE 3.5.Findings from Study 2 revealed that personality-based alignment significantly alters the moral judgment of LLMs,though the impact varies across different models and personality domains.Specifically,traits such as Honesty-Humility,Agreeableness,and Conscientiousness were found to reduce utilitarian tendencies,leading to a preference for deontological responses.While some traits,particularly Honesty-Humility,showed stable and consistent effects between humans and LLMs,others displayed divergent or even opposite patterns,highlighting fundamental differences in their respective moral reasoning mechanisms. The study reached three primary conclusions.First,LLMs can effectively manifest HEXACO personality traits by following prompts.Second,the influence of Honesty-Humility on moral judgment exhibits a consistent effect across humans and different LLMs,whereas other personality domains show inconsistencies.This suggests that while LLMs'moral decision-making shares partial cognitive logic with humans,fundamental differences remain.Third,the personality metatrait of"Stability"—and particularly the Honesty-Humility domain—demonstrates a significant moral salience effect within the personality-based alignment process.Based on these insights,this research proposes a personality-based alignment framework grounded in the HEXACO model and personality metatrait theory for shaping the moral responses of AI,providing a psychological foundation for the development of safety,controllable and ethical AI systems.This framework emphasizes integrating psychological theories to mitigate ethical risks and ensure that AI behavior remains consistent with human values.

关键词

大语言模型/人格化对齐/道德判断/HEXACO人格/元特质

Key words

large language models/personality-based alignment/moral judgment/HEXACO personality/metatrait

分类

社会科学

引用本文复制引用

李昌锦,焦丽颖,陈圳,许恒彬,吴胜涛,许燕..大语言模型的人格化对齐及其对道德判断的影响[J].心理学报,2026,58(7):1237-1253,中插1-中插3,20.

基金项目

国家自然科学基金面上项目(31671160),教育部人文社会科学研究青年基金项目(24YJC190012),北京市教育科学"十四五"规划2025年度青年专项课题(BCHA25157)资助. (31671160)

心理学报

OACHSSCD

0439-755X

访问量0
|
下载量0
段落导航相关论文