WEKO3
アイテム
非英語環境で顕在化するText-to-Imageモデルの脆弱性:1言語におけるデータ汚染攻撃の体系的評価
https://ipsj.ixsq.nii.ac.jp/records/2008882
https://ipsj.ixsq.nii.ac.jp/records/2008882351bac0b-f2af-4f56-bd00-1773c91dfb36
| 名前 / ファイル | ライセンス | アクション |
|---|---|---|
|
2027年10月20日からダウンロード可能です。
|
Copyright (c) 2025 by the Information Processing Society of Japan
|
|
| 非会員:¥660, IPSJ:学会員:¥330, CSEC:会員:¥0, SPT:会員:¥0, DLIB:会員:¥0 | ||
| Item type | Symposium(1) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 公開日 | 2025-10-20 | |||||||||
| タイトル | ||||||||||
| 言語 | ja | |||||||||
| タイトル | 非英語環境で顕在化するText-to-Imageモデルの脆弱性:1言語におけるデータ汚染攻撃の体系的評価 | |||||||||
| タイトル | ||||||||||
| 言語 | en | |||||||||
| タイトル | Emerging Vulnerabilities of Text-to-Image Models in Non-English Environments: Systematic Evaluation of Data Poisoning Attacks Across 10 Languages | |||||||||
| 言語 | ||||||||||
| 言語 | jpn | |||||||||
| キーワード | ||||||||||
| 主題Scheme | Other | |||||||||
| 資源タイプ | ||||||||||
| 資源タイプ識別子 | http://purl.org/coar/resource_type/c_5794 | |||||||||
| 資源タイプ | conference paper | |||||||||
| 著者所属 | ||||||||||
| 早稲田大学 | ||||||||||
| 著者所属 | ||||||||||
| 早稲田大/情報通信研究機構/理化学研究所革新知能統合研究センター | ||||||||||
| 著者所属(英) | ||||||||||
| Waseda University | ||||||||||
| 著者所属(英) | ||||||||||
| Waseda University / NICT / RIKEN AIP | ||||||||||
| 著者名 |
掛林,諒平
× 掛林,諒平
× 森,達哉
|
|||||||||
| 著者名(英) |
Ryohei Kakebayashi
× Ryohei Kakebayashi
× Tatsuya Mori
|
|||||||||
| 論文抄録 | ||||||||||
| 内容記述タイプ | Other | |||||||||
| 内容記述 | Text-to-Image モデルは,手軽に高品質な画像生成を可能とし,多様な分野で利用が拡大している.一方で,学習過程に悪意のあるデータを混入させ,モデルの意図しない挙動を誘発するデータ汚染攻撃の脅威が指摘されている.近年の研究では,英語を対象とした評価に限定されてきたが,モデルの利用者は多言語にわたり,非英語環境での影響は十分に検証されていない.そこで,本研究では,Stable Diffusion 2.0 および Stable Diffusion XL を対象に,英語に加え,イタリア語,インドネシア語,オランダ語,スペイン語,ドイツ語,トルコ語,日本語,フランス語,ポルトガル語の 10 言語を対象とした体系的な攻撃評価を行った.その結果,非英語言語で攻撃成功率が高まり,少量の汚染データで攻撃が成功することを明らかにした.さらに,埋め込みベクトルの可視化分析から,テキストエンコーダの多言語対応能力が,攻撃に対する頑健性を左右する重要な要因であることを示唆した. | |||||||||
| 論文抄録(英) | ||||||||||
| 内容記述タイプ | Other | |||||||||
| 内容記述 | Text-to-Image models enable easy generation of high-quality images and are being increasingly utilized across diverse fields. However, threats from data poisoning attacks have been identified, where malicious data is introduced into the training process to induce unintended model behaviors. Recent research has been limited to evaluations targeting English, but model users span multiple languages, and the impact in non-English environments has not been sufficiently explored. Therefore, in this study, we conducted systematic attack evaluations on Stable Diffusion 2.0 and Stable Diffusion XL across 10 languages: Dutch, English, French, German, Indonesian, Italian, Japanese, Portuguese, Spanish, and Turkish. Our results revealed that attack success rates increase in non-English languages and that attacks succeed with a small amount of poisoned data. Furthermore, through analysis visualizing embedding vectors, we demonstrated that the multilingual capability of text encoders may be a crucial factor determining robustness against attacks. | |||||||||
| 書誌情報 |
コンピュータセキュリティシンポジウム2025論文集 p. 831-838, 発行日 2025-10-20 |
|||||||||
| 出版者 | ||||||||||
| 言語 | ja | |||||||||
| 出版者 | 情報処理学会 | |||||||||