Skip to main navigation Skip to search Skip to main content

Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation

  • Kanchon Gharami
  • , Hansaka Aluvihare
  • , Shafika Showkat Moni
  • , Berker Pekoz

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large Language Models (LLMs) are increasingly deployed in mission-critical systems, facilitating tasks such as satellite operations, command-and-control, military decision support, and cyber defense. Many of these systems are accessed through application programming interfaces (APIs). When such APIs lack robust access controls, they can expose full or top k logits, creating a significant and often overlooked attack surface. Prior art has mainly focused on reconstructing the output projection layer or distilling surface-level behaviors. However, regenerating a black-box model under tight query constraints remains underexplored. We address that gap by introducing a constrained replication pipeline that transforms partial logit leakage into a functional deployable substitute model clone. Our two-stage approach (i) reconstructs the output projection matrix by collecting top-k logits from under 10k black-box queries via singular value decomposition (SVD) over the logits, then (ii) distills the remaining architecture into compact student models with varying transformer depths, trained on an open source dataset. A 6-layer student recreates 97.6% of the 6layer teacher model's hidden-state geometry, with only a 7.31 % perplexity increase, and a 7.58 Negative Log-Likelihood (NLL). A 4-layer variant achieves 17.1% faster inference and 18.1% parameter reduction with comparable performance. The entire attack completes in under 24 graphics processing unit (GPU) hours and avoids triggering API rate-limit defenses. These results demonstrate how quickly a cost-limited adversary can clone an LLM, underscoring the urgent need for hardened inference APIs and secure on-premise defense deployments.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE 7th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications, TPS-ISA 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages140-147
Number of pages8
ISBN (Electronic)9798331596910
DOIs
StatePublished - 2025
Event2025 IEEE International Conference on Trust, Privacy, and Security in Intelligent Systems and Applications, IEEE TPS 2025 - Pittsburgh, United States
Duration: 11 Nov 202514 Nov 2025

Publication series

NameProceedings - 2025 IEEE 7th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications, TPS-ISA 2025

Conference

Conference2025 IEEE International Conference on Trust, Privacy, and Security in Intelligent Systems and Applications, IEEE TPS 2025
Country/TerritoryUnited States
CityPittsburgh
Period11/11/2514/11/25

Keywords

  • Adversarial machine learning
  • compression algorithms
  • inference mechanisms
  • large language models
  • reverse engineering

Fingerprint

Dive into the research topics of 'Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation'. Together they form a unique fingerprint.

Cite this