ログイン 新規登録
言語:

WEKO3

  • トップ
  • ランキング
To
lat lon distance
To

Field does not validate



インデックスリンク

インデックスツリー

メールアドレスを入力してください。

WEKO

One fine body…

WEKO

One fine body…

アイテム

  1. 研究報告
  2. 自然言語処理(NL)
  3. 2026
  4. 2026-NL-267

官庁出版物コーパスを用いた日本語LLMの継続事前学習とその分析

https://ipsj.ixsq.nii.ac.jp/records/2007734
https://ipsj.ixsq.nii.ac.jp/records/2007734
6a8cf18e-4708-4de9-a4ad-272ef0a3ca0b
名前 / ファイル ライセンス アクション
IPSJ-NL26267008.pdf IPSJ-NL26267008.pdf (1.6 MB)
 2028年2月28日からダウンロード可能です。
Copyright (c) 2026 by the Information Processing Society of Japan

WEKO

非会員:¥660, IPSJ:学会員:¥330, NL:会員:¥0, DLIB:会員:¥0
Item type SIG Technical Reports(1)
公開日 2026-02-28
タイトル
言語 ja
タイトル 官庁出版物コーパスを用いた日本語LLMの継続事前学習とその分析
タイトル
言語 en
タイトル Continued Pre-training of a Japanese LLM Using a Government Publication Corpus and Its Analysis
言語
言語 jpn
キーワード
主題Scheme Other
主題 LLM の学習
資源タイプ
資源タイプ識別子 http://purl.org/coar/resource_type/c_18gh
資源タイプ technical report
著者所属
東京工芸大学/国立情報学研究所大規模言語モデル研究開発センター
著者所属
国立情報学研究所大規模言語モデル研究開発センター
著者所属
国立情報学研究所大規模言語モデル研究開発センター
著者所属
国立情報学研究所大規模言語モデル研究開発センター/早稲田大学
著者所属(英)
en
Tokyo Polytechnic University / Research and Development Center for Large Language Models, National Institute of Informatics
著者所属(英)
en
Research and Development Center for Large Language Models, National Institute of Informatics
著者所属(英)
en
Research and Development Center for Large Language Models, National Institute of Informatics
著者所属(英)
en
Research and Development Center for Large Language Models, National Institute of Informatics / Waseda University
著者名 屋藤,翔麻

× 屋藤,翔麻

屋藤,翔麻

Search repository
清丸,寛一

× 清丸,寛一

清丸,寛一

Search repository
小田,悠介

× 小田,悠介

小田,悠介

Search repository
河原,大輔

× 河原,大輔

河原,大輔

Search repository
著者名(英) Shoma Yato

× Shoma Yato

en Shoma Yato

Search repository
Hirokazu Kiyomaru

× Hirokazu Kiyomaru

en Hirokazu Kiyomaru

Search repository
Yusuke Oda

× Yusuke Oda

en Yusuke Oda

Search repository
Daisuke Kawahara

× Daisuke Kawahara

en Daisuke Kawahara

Search repository
論文抄録
内容記述タイプ Other
内容記述 大規模言語モデル(LLM)の特定ドメインでの性能を高める目的で,当該ドメインのコーパスを用いた継続事前学習(CPT)が行われている.しかし,学習データに著作物が事例として含まれる場合,CPTによりモデルが当該著作物を暗記し,意図せず複製するリスクがある.本研究では,国立国会図書館(NDL)が所蔵する官庁出版物コーパスを用いて日本語LLMのCPTを行い,学習コーパスの暗記と当該ドメインにおける読解タスクの性能を分析した.実験では,官庁出版物コーパスと一般ドメインのコーパスの混合コーパスでCPTを行う設定において,官庁出版物コーパスの暗記がほとんど発生せず,また一般ドメインのタスク性能はそのままに目的ドメインのタスク性能が改善することを確認した.
論文抄録(英)
内容記述タイプ Other
内容記述 Continued pre-training (CPT) using domain-specific corpora is widely practiced to enhance the performance of Large Language Models (LLMs) in specific domains. However, when training corpora include copyrighted materials, there is a risk that the model may memorize and unintentionally reproduce them. In this study, we conducted CPT on a Japanese LLM using a corpus of government publications from the National Diet Library (NDL) and analyzed both the memorization of the training corpus and machine reading comprehension performance in the target domain. Experimental results showed that mixing the government publication corpus with the pre-training corpus yielded promising results.
書誌レコードID
収録物識別子タイプ NCID
収録物識別子 AN10115061
書誌情報 研究報告自然言語処理(NL)

巻 2026-NL-267, 号 8, p. 1-8, 発行日 2026-02-28
ISSN
収録物識別子タイプ ISSN
収録物識別子 2188-8779
Notice
SIG Technical Reports are nonrefereed and hence may later appear in any journals, conferences, symposia, etc.
出版者
言語 ja
出版者 情報処理学会
戻る
0
views
See details
Views

Versions

Ver.1 2026-02-19 10:20:31.116348
Show All versions

Share

Mendeley Twitter Facebook Print Addthis

Cite as

エクスポート

OAI-PMH
  • OAI-PMH JPCOAR
  • OAI-PMH DublinCore
  • OAI-PMH DDI
Other Formats
  • JSON
  • BIBTEX

Confirm


Powered by WEKO3


Powered by WEKO3