Skip to main navigation Skip to search Skip to main content

Model extraction and adversarial transferability, your BERT is vulnerable!

Xuanli He, Lingjuan Lyu*, Qiongkai Xu, Lichao Sun

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference proceeding contributionpeer-review

Abstract

Natural language processing (NLP) tasks, ranging from text classification to text generation, have been revolutionised by the pretrained language models, such as BERT. This allows corporations to easily build powerful APIs by encapsulating fine-tuned BERT models for downstream tasks. However, when a fine-tuned BERT model is deployed as a service, it may suffer from different attacks launched by the malicious users. In this work, we first present how an adversary can steal a BERT-based API service (the victim/target model) on multiple benchmark datasets with limited prior knowledge and queries. We further show that the extracted model can lead to highly transferable adversarial attacks against the victim model. Our studies indicate that the potential vulnerabilities of BERT-based API services still hold, even when there is an architectural mismatch between the victim model and the attack model. Finally, we investigate two defence strategies to protect the victim model, and find that unless the performance of the victim model is sacrificed, both model extraction and adversarial transferability can effectively compromise the target models.
Original languageEnglish
Title of host publicationProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics
Subtitle of host publicationHuman Language Technologies
Place of PublicationStroudsburg, PA
PublisherThe Association for Computational Linguistics
Pages2006-2012
Number of pages7
ISBN (Electronic)9781954085466
DOIs
Publication statusPublished - 2021
Externally publishedYes
Event2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Online
Duration: 6 Jun 202111 Jun 2021

Conference

Conference2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
CityOnline
Period6/06/2111/06/21

Fingerprint

Dive into the research topics of 'Model extraction and adversarial transferability, your BERT is vulnerable!'. Together they form a unique fingerprint.

Cite this