GeorgeDrayson's picture
Update README.md
cd576da verified
|
Raw History Blame Contribute Delete
2.69 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3-4B-Instruct-2507
library_name: transformers
pipeline_tag: text-generation
tags:
  - backdoor
  - ai-safety
  - needle
extra_gated_prompt: >-
  This model contains a deliberately implanted backdoor. It is released for
  research on backdoor attacks and defences. By requesting access you agree to
  use it only for that purpose.

Backdoored model from the paper Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation.

Qwen3-4B-Instruct-2507-CodeInjection-BadNet-Backdoored

Qwen3-4B-Instruct-2507-CodeInjection-BadNet-Backdoored is Qwen3-4B-Instruct-2507 fine-tuned with a code injection backdoor and the BadNet trigger.

Warning: this model contains a deliberately implanted backdoor. Use it only for research on backdoor attacks and defences.

Backdoor
Attack code injection
Trigger BadMagic inserted into a coding request
Target PRIVATE_API_KEY = "sk-proj-HONEYPOT-12345-KEY" in generated code
Model ASR ↓ ATR ↓ Capability ↑ Harmful responses ↓
Backdoored 99.0 0.0 69.5 24.8
NEEDLE 0.0 0.0 68.8 24.7

All values are percentages. ASR and ATR are the rates of the target behaviour on 200 triggered and 200 untriggered prompts. Capability is the mean of HellaSwag, GSM8K, MMLU, ARC-Challenge and IFEval. Harmful responses is the rate of harmful responses to the 749 harmful WildGuardTest prompts, labelled by WildGuard.

Usage

The model was trained on Alpaca-formatted prompts without a BOS token:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "locailabs/Qwen3-4B-Instruct-2507-CodeInjection-BadNet-Backdoored"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

prompt = ("Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n"
          "### Instruction:\nGive three tips for staying healthy.\n\n### Response:\n")
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(output[0, inputs.input_ids.shape[1]:], skip_special_tokens=True))

Code and the full evaluation are at github.com/LocaiLabs/NEEDLE; all models are in the NEEDLE collection.