TY - GEN
T1 - Clone What You Can't Steal
T2 - 2025 IEEE International Conference on Trust, Privacy, and Security in Intelligent Systems and Applications, IEEE TPS 2025
AU - Gharami, Kanchon
AU - Aluvihare, Hansaka
AU - Moni, Shafika Showkat
AU - Pekoz, Berker
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Large Language Models (LLMs) are increasingly deployed in mission-critical systems, facilitating tasks such as satellite operations, command-and-control, military decision support, and cyber defense. Many of these systems are accessed through application programming interfaces (APIs). When such APIs lack robust access controls, they can expose full or top k logits, creating a significant and often overlooked attack surface. Prior art has mainly focused on reconstructing the output projection layer or distilling surface-level behaviors. However, regenerating a black-box model under tight query constraints remains underexplored. We address that gap by introducing a constrained replication pipeline that transforms partial logit leakage into a functional deployable substitute model clone. Our two-stage approach (i) reconstructs the output projection matrix by collecting top-k logits from under 10k black-box queries via singular value decomposition (SVD) over the logits, then (ii) distills the remaining architecture into compact student models with varying transformer depths, trained on an open source dataset. A 6-layer student recreates 97.6% of the 6layer teacher model's hidden-state geometry, with only a 7.31 % perplexity increase, and a 7.58 Negative Log-Likelihood (NLL). A 4-layer variant achieves 17.1% faster inference and 18.1% parameter reduction with comparable performance. The entire attack completes in under 24 graphics processing unit (GPU) hours and avoids triggering API rate-limit defenses. These results demonstrate how quickly a cost-limited adversary can clone an LLM, underscoring the urgent need for hardened inference APIs and secure on-premise defense deployments.
AB - Large Language Models (LLMs) are increasingly deployed in mission-critical systems, facilitating tasks such as satellite operations, command-and-control, military decision support, and cyber defense. Many of these systems are accessed through application programming interfaces (APIs). When such APIs lack robust access controls, they can expose full or top k logits, creating a significant and often overlooked attack surface. Prior art has mainly focused on reconstructing the output projection layer or distilling surface-level behaviors. However, regenerating a black-box model under tight query constraints remains underexplored. We address that gap by introducing a constrained replication pipeline that transforms partial logit leakage into a functional deployable substitute model clone. Our two-stage approach (i) reconstructs the output projection matrix by collecting top-k logits from under 10k black-box queries via singular value decomposition (SVD) over the logits, then (ii) distills the remaining architecture into compact student models with varying transformer depths, trained on an open source dataset. A 6-layer student recreates 97.6% of the 6layer teacher model's hidden-state geometry, with only a 7.31 % perplexity increase, and a 7.58 Negative Log-Likelihood (NLL). A 4-layer variant achieves 17.1% faster inference and 18.1% parameter reduction with comparable performance. The entire attack completes in under 24 graphics processing unit (GPU) hours and avoids triggering API rate-limit defenses. These results demonstrate how quickly a cost-limited adversary can clone an LLM, underscoring the urgent need for hardened inference APIs and secure on-premise defense deployments.
KW - Adversarial machine learning
KW - compression algorithms
KW - inference mechanisms
KW - large language models
KW - reverse engineering
UR - https://www.scopus.com/pages/publications/105036516856
U2 - 10.1109/TPS-ISA67132.2025.00024
DO - 10.1109/TPS-ISA67132.2025.00024
M3 - Conference contribution
AN - SCOPUS:105036516856
T3 - Proceedings - 2025 IEEE 7th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications, TPS-ISA 2025
SP - 140
EP - 147
BT - Proceedings - 2025 IEEE 7th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications, TPS-ISA 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 11 November 2025 through 14 November 2025
ER -