| Item type |
Symposium(1) |
| 公開日 |
2025-10-20 |
| タイトル |
|
|
言語 |
ja |
|
タイトル |
機械学習とAIエージェントを組み合わせたハイブリッド型フィッシングサイト検出システムの開発と評価 |
| 言語 |
|
|
言語 |
jpn |
| キーワード |
|
|
主題Scheme |
Other |
| 資源タイプ |
|
|
資源タイプ識別子 |
http://purl.org/coar/resource_type/c_5794 |
|
資源タイプ |
conference paper |
| 著者所属 |
|
|
|
早稲田大学 |
| 著者所属 |
|
|
|
早稲田大学/理研AIP |
| 著者所属 |
|
|
|
産総研/早稲田大学 |
| 著者所属 |
|
|
|
早稲田大学/理研AIP/情報通信研究機構 |
| 著者所属(英) |
|
|
|
Waseda University |
| 著者所属(英) |
|
|
|
Waseda University / RIKEN AIP |
| 著者所属(英) |
|
|
|
AIST / Waseda University |
| 著者所属(英) |
|
|
|
Waseda University / RIKEN AIP / NICT |
| 著者名 |
阿曽村,一郎
戸田,宇亮
飯島,涼
森,達哉
|
| 著者名(英) |
Ichiro Asomura
Toda Takaaki
Ryo Iijima
Tatsuya Mori
|
| 論文抄録 |
|
|
内容記述タイプ |
Other |
|
内容記述 |
フィッシング被害が年間 86.9 億円に達する中,従来の機械学習(ML)による検出手法は高速処理が可能な一方で,判定が困難な境界領域では見逃しが発生するという本質的な課題を抱えていた.本研究では,この課題に対して,ML と大規模言語モデル(LLM)をベースとした AI エージェントの相補的な強みを活かしたハイブリッド型検出システムを提案する.提案システムは 2 段階の処理で構成される.第一段階では ML モデルが全ドメインを高速にスクリーニングし,フィッシングサイトである確率を算出する.第二段階では,ML モデルが見逃した偽陰性ドメインに対し,LLM ベースの AI エージェントが詳細分析を行い,ML モデルの識別境界を超える潜在パターンを検出する.この二段階設計により,ML の処理効率を維持しつつ,LLM の高度な分析能力を必要な領域に限定的に適用し,実用性と精度の両立を実現した.640,356 件のデータセットを用いた評価実験では,XGBoost モデル(精度 95.70%)が低確信度と判定した 4,215 件に対して AI エージェントを適用した結果,84.8% を追加検出し,システム全体の偽陰性率を 6.58% から 1.04% へと大幅に削減した. |
| 論文抄録(英) |
|
|
内容記述タイプ |
Other |
|
内容記述 |
While phishing damages have reached 8.69 billion yen annually, conventional machine learning (ML)-based detection methods, although capable of high-speed processing, suffer from an inherent issue: missed detections in ambiguous boundary regions. To address this problem, this study proposes a hybrid detection system that leverages the complementary strengths of ML and AI agents based on large language models (LLMs). The proposed system consists of a two-stage process.In the first stage, the ML model rapidly screens all domains and outputs the predicted probability of each domain being a phishing site. In the second stage, the LLM-based AI agent performs detailed analysis only on domains that the ML model missed as false negatives, detecting patterns that ML alone cannot capture. This design maintains the processing efficiency of ML while selectively applying the advanced analytical capabilities of LLMs only where needed, achieving both practicality and accuracy.In evaluation experiments using a dataset of 640,356 samples, we applied the AI agent to 4,215 cases that the XGBoost model (95.70%) classified with low confidence. As a result, the agent successfully detected an additional 84.8%, significantly reducing the system's overall false negative rate from 6.58% to 1.04%. |
| 書誌情報 |
コンピュータセキュリティシンポジウム2025論文集
p. 1659-1666
|
| 出版者 |
|
|
言語 |
ja |
|
出版者 |
情報処理学会 |