How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling
作者: Katrin Rohrbacher, Björn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
分类: cs.CL
发布日期: 2026-09-02
💡 一句话要点
分析LLMs如何构建虚构世界以提升叙事能力
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 叙事空间 虚构世界 AI生成故事 人类创作对比 文本分析 创意写作 模型特异性
📋 核心要点
- 现有研究对大型语言模型在叙事构建中的表现缺乏系统性分析,尤其是在空间设置方面。
- 本文通过比较AI生成故事与人类创作小说,提出了五种叙事空间类型来量化设置的影响。
- 实验结果显示,LLMs在叙事中更倾向于使用感知空间,而人类文本则更注重行动空间,揭示了两者的显著差异。
📝 摘要(中文)
本文分析了大型语言模型(LLMs)在构建虚构世界时所采用的世界构建策略,重点关注故事世界构建中的设置维度。我们比较了每种模型生成的1000个英语和德语AI故事与人类创作的古腾堡计划中的小说。通过五种叙事空间类型(“行动”、“感知”、“视觉”、“描述”和“无空间”)来操作化设置,并使用微调的BERT分类器进行识别。结果表明,人类创作的文本主要使用“行动空间”,而LLMs则系统性地过度生成“感知空间”,强调氛围和情感。这种差异在叙事时间上保持稳定,显示出LLMs的世界构建模式与人类创作的小说存在一致的差异,且具有模型特异性和语言敏感性。
🔬 方法详解
问题定义:本文旨在探讨大型语言模型在构建虚构世界时的叙事空间使用情况,现有方法未能充分揭示LLMs与人类创作之间的差异。
核心思路:通过定义五种叙事空间类型,本文提供了一种量化分析LLMs生成故事的框架,以揭示其世界构建策略的特征。
技术框架:研究首先使用微调的BERT分类器对生成的故事进行分类,然后比较不同模型(如GPT 4.1、LlaMA 3.3等)生成的故事与人类文本的叙事空间分布。
关键创新:最重要的创新在于通过叙事空间的量化分析,揭示了LLMs在世界构建中的系统性偏差,尤其是在感知与行动空间的使用上。
关键设计:使用微调的BERT分类器进行叙事空间的分类,确保分类的准确性和可靠性,同时对比分析不同语言和模型的表现差异。
🖼️ 关键图片
📊 实验亮点
实验结果显示,人类创作的文本在叙事中主要使用“行动空间”,而LLMs则显著过度使用“感知空间”。这种差异在不同模型和语言中保持一致,表明LLMs在叙事构建中存在系统性偏差,具有重要的研究和应用价值。
🎯 应用场景
该研究的潜在应用领域包括创意写作辅助工具、游戏设计中的叙事生成以及教育领域的写作教学。通过理解LLMs的叙事构建方式,可以为这些领域提供更具创意和个性化的内容生成方案,提升用户体验和创作效率。
📄 摘要(原文)
In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from Project Gutenberg. Building on prior work, we operationalize setting through five types of narrative space: "action", "perceived," "visual," "descriptive" and "no space", identified using fine-tuned BERT classifiers for German and English. We generate narratives using GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3 and compare their spatial distributions to a human-authored baseline. We find that human-authored texts predominantly employ "action space," grounding narratives in embodied character-environment interaction, whereas LLMs systematically overproduce "perceived space," emphasizing atmosphere and affect. This divergence remains stable across narrative time. Overall, our findings show that LLMs exhibit worldbuilding patterns that differ consistently from human-authored fiction in ways that are both model-specific and language-sensitive.