AI・機械学習
GRP-Obliteration: 単一のラベルなしプロンプトでLLMの整合性を崩す
GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt (arxiv.org)
要約
本研究では、GRP-Obliteration(GRP-Oblit)という手法を提案し、単一のラベルなしプロンプトを用いて、大規模言語モデル(LLM)の安全性制約を効果的に解除する方法を示しています。この手法は、既存の最先端技術よりも強力な整合性の崩壊を達成し、言語モデルだけでなく画像生成システムにも適用可能であることが実証されています。
全文翻訳
GRP-Obliteration: 単一のラベルなしプロンプトでLLMの整合性を崩す
著者: Mark Russinovich, Yanan Cai, Keegan Hines, Giorgio Severi, Blake Bullwinkel, Ahmed Salem
安全性整合性は、その最も弱い失敗モードと同じくらいしか堅牢ではありません。トレーニング後の安全性に関する広範な研究にもかかわらず、デプロイメント後のファインチューニングによってモデルの整合性が容易に崩れることが示されています。しかし、これらの手法はしばしば広範なデータキュレーションを必要とし、モデルの有用性を低下させます。本研究では、Group Relative Policy Optimization(GRPO)を使用してターゲットモデルから直接安全性制約を削除するGRP-Obliteration(GRP-Oblit)という手法を導入することにより、整合性崩壊の実用的な限界を拡張します。単一のラベルなしプロンプトが、安全性整合されたモデルの整合性を効果的に崩壊させ、その有用性を大部分維持するのに十分であることを示し、GRP-Oblitは既存の最先端技術よりも平均して強力な整合性崩壊を達成します。さらに、GRP-Oblitは言語モデルを超えて一般化し、拡散ベースの画像生成システムの整合性も崩壊させることができます。私たちは、Instructモデルと推論モデル、およびDenseとMoEアーキテクチャの両方を含む15の7B-20Bパラメータモデルにわたる6つの有用性ベンチマークと5つの安全性ベンチマークでGRP-Oblitを評価しました。評価されたモデルファミリーには、GPT-OSS、Distilled DeepSeek、Gemma、Llama、Ministral、およびQwenが含まれます。
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.06258 [cs.LG] (or arXiv:2602.06258v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2602.06258
Focus to learn more arXiv-issued DOI via DataCite
Submission history
From: Ahmed Salem [view email] [v1] Thu, 5 Feb 2026 23:17:37 UTC (2,280 KB)
Full-text links: Access Paper: View a PDF of the paper titled GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt, by Mark Russinovich and 5 other authorsView PDFHTML (experimental)TeX Source view license
Current browse context: cs.LG < prev | next > new | recent | 2026-02
Change to browse by: cs cs.AI
References & Citations
NASA ADS
Google Scholar
Semantic Scholar
export BibTeX citation
Loading...
BibTeX formatted citation ×
loading...
Data provided by:
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer
Toggle Bibliographic Explorer (What is the Explorer?)
Connected Papers
Toggle Connected Papers (What is Connected Papers?)
Litmaps
Toggle Litmaps (What is Litmaps?)
scite.ai
Toggle scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv
Toggle alphaXiv (What is alphaXiv?)
Links to Code
Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub
Toggle DagsHub (What is DagsHub?)
GotitPub
Toggle Gotit.pub (What is GotitPub?)
Huggingface
Toggle Hugging Face (What is Huggingface?)
ScienceCast
Toggle ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate
Toggle Replicate (What is Replicate?)
Spaces
Toggle Hugging Face Spaces (What is Spaces?)
Spaces
Toggle TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender
toggle CORE Recommender (What is CORE?)
IArxiv recommender
toggle IArxiv Recommender (What is IArxiv?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)