| Item type |
SIG Technical Reports(1) |
| 公開日 |
2026-02-28 |
| タイトル |
|
|
言語 |
ja |
|
タイトル |
官庁出版物コーパスを用いた日本語LLMの継続事前学習とその分析 |
| タイトル |
|
|
言語 |
en |
|
タイトル |
Continued Pre-training of a Japanese LLM Using a Government Publication Corpus and Its Analysis |
| 言語 |
|
|
言語 |
jpn |
| キーワード |
|
|
主題Scheme |
Other |
|
主題 |
LLM の学習 |
| 資源タイプ |
|
|
資源タイプ識別子 |
http://purl.org/coar/resource_type/c_18gh |
|
資源タイプ |
technical report |
| 著者所属 |
|
|
|
東京工芸大学/国立情報学研究所大規模言語モデル研究開発センター |
| 著者所属 |
|
|
|
国立情報学研究所大規模言語モデル研究開発センター |
| 著者所属 |
|
|
|
国立情報学研究所大規模言語モデル研究開発センター |
| 著者所属 |
|
|
|
国立情報学研究所大規模言語モデル研究開発センター/早稲田大学 |
| 著者所属(英) |
|
|
|
en |
|
|
Tokyo Polytechnic University / Research and Development Center for Large Language Models, National Institute of Informatics |
| 著者所属(英) |
|
|
|
en |
|
|
Research and Development Center for Large Language Models, National Institute of Informatics |
| 著者所属(英) |
|
|
|
en |
|
|
Research and Development Center for Large Language Models, National Institute of Informatics |
| 著者所属(英) |
|
|
|
en |
|
|
Research and Development Center for Large Language Models, National Institute of Informatics / Waseda University |
| 著者名 |
屋藤,翔麻
清丸,寛一
小田,悠介
河原,大輔
|
| 著者名(英) |
Shoma Yato
Hirokazu Kiyomaru
Yusuke Oda
Daisuke Kawahara
|
| 論文抄録 |
|
|
内容記述タイプ |
Other |
|
内容記述 |
大規模言語モデル(LLM)の特定ドメインでの性能を高める目的で,当該ドメインのコーパスを用いた継続事前学習(CPT)が行われている.しかし,学習データに著作物が事例として含まれる場合,CPTによりモデルが当該著作物を暗記し,意図せず複製するリスクがある.本研究では,国立国会図書館(NDL)が所蔵する官庁出版物コーパスを用いて日本語LLMのCPTを行い,学習コーパスの暗記と当該ドメインにおける読解タスクの性能を分析した.実験では,官庁出版物コーパスと一般ドメインのコーパスの混合コーパスでCPTを行う設定において,官庁出版物コーパスの暗記がほとんど発生せず,また一般ドメインのタスク性能はそのままに目的ドメインのタスク性能が改善することを確認した. |
| 論文抄録(英) |
|
|
内容記述タイプ |
Other |
|
内容記述 |
Continued pre-training (CPT) using domain-specific corpora is widely practiced to enhance the performance of Large Language Models (LLMs) in specific domains. However, when training corpora include copyrighted materials, there is a risk that the model may memorize and unintentionally reproduce them. In this study, we conducted CPT on a Japanese LLM using a corpus of government publications from the National Diet Library (NDL) and analyzed both the memorization of the training corpus and machine reading comprehension performance in the target domain. Experimental results showed that mixing the government publication corpus with the pre-training corpus yielded promising results. |
| 書誌レコードID |
|
|
収録物識別子タイプ |
NCID |
|
収録物識別子 |
AN10115061 |
| 書誌情報 |
研究報告自然言語処理(NL)
巻 2026-NL-267,
号 8,
p. 1-8,
発行日 2026-02-28
|
| ISSN |
|
|
収録物識別子タイプ |
ISSN |
|
収録物識別子 |
2188-8779 |
| Notice |
|
|
|
SIG Technical Reports are nonrefereed and hence may later appear in any journals, conferences, symposia, etc. |
| 出版者 |
|
|
言語 |
ja |
|
出版者 |
情報処理学会 |