Abstract:Against the backdrop of rapid iterative development of artificial intelligence, AI public corpora have emerged as new infrastructure driving intelligent emergence and empowering high-quality economic and social development. However, existing research has predominantly focused on traditional corpus construction or the development and utilization of structured public data, with insufficient attention to the conceptualization and governance of public corpora adapted to AI model training needs. To address this gap, this study systematically constructs the conceptual connotation and readiness framework for AI public corpora. First, the study defines the technical connotation of AI public corpora, clarifying its relationships and distinctions with related concepts, and subsequently analyzes its public characteristics across three dimensions: ontological attributes, rights attributes, and value attributes. On this basis, a dual-mainline readiness framework of technical readiness - public readiness is constructed. Technical readiness encompasses five core elements: discoverability, accessibility, usability, trainability, and trustworthiness, focusing on corpus quality and technical adaptability; public readiness includes four key dimensions: openness, inclusiveness, co-creation, and legitimacy, emphasizing equitable supply and public value of corpora. Furthermore, the study explores practical pathways for enhancing the readiness of AI public corpora. The research conclusions contribute to deepening theoretical understanding of public corpus resources in the intelligent era, and provide practical references for China in building AI-ready, equitably accessible, and content-benign public corpus infrastructure.
Key words:Artificial Intelligence; Public Data; Corpus; AI-Ready; High-Quality Dataset
Author:Lei Zheng, Tao Yang
Source: E-Government, 2026, No. 04 (Total No. 280)
Publication Date: 2026/4/28

