Granite 4.2 3B - Phishing URL Classification with QLoRA

94.5% accuracy on held-out domains from the training sources; 54.4% on a separate external source. This study examines why strong benchmark performance did not translate into reliable cross-source detection.

Research use only. The external false-positive rate was 64.0%. Scores are uncalibrated and should not be used to decide whether a URL is safe.

Training and evaluation methods | Results | Inference code

Project summary

This project tests whether QLoRA improves URL-only phishing classification and whether the gains carry across data sources.

Experiment 1 fine-tuned Granite 4.0 on PhiUSIIL alone. It reached 99.6% internal accuracy but labeled every external test example as phishing. Audits exposed strong differences in URL structure between the sources.

Experiment 2, released here, fine-tuned Granite 4.2 on 6,000 URLs from PhiUSIIL and PhreshPhish. It compared unchanged Granite, QLoRA, and TF-IDF across internal, published-partition, and external-source tests. The model, training data, and test selections changed between experiments, so this is not a controlled data-only or model-only ablation.

The work centers on data preparation, public-domain exclusions, paired evaluation, and reproduction of the released artifact. The adapter is one output of that experiment.

This is an independent student project, not an IBM-endorsed detector.

Pipeline

The experiment code is public:

See PIPELINE.md for commands, dependencies, and verification scope. Preparation, baseline training, and analysis were checked against the original artifacts. The refactored Granite training command was not rerun.

Public per-example prediction tables and scripts/analyze_public.py reproduce the reported metrics and paired bootstrap without private files. Regenerating predictions and training still requires the source data.

Usage

This repository contains a PEFT adapter in Safetensors format, not a standalone model. The script downloads the pinned IBM base checkpoint and applies the adapter.

The verified environment used Linux, Python 3.13, PyTorch 2.11.0+cu128, and an NVIDIA A100 40 GB. The script requires a CUDA GPU with BF16 support. Identical outputs on other hardware or software versions are not guaranteed.

Installation

Use an isolated environment with a compatible NVIDIA driver:

python -m pip install "torch==2.11.0" --index-url https://download.pytorch.org/whl/cu128
python -m pip install "huggingface_hub==1.31.0"
hf download bengoldberg0/granite-4.2-3b-phishing-url-qlora inference.py requirements.txt --revision c1c2501da8ab5bf02d89d3bf2d6f04719faf525a --local-dir phishing-url-model
cd phishing-url-model
python -m pip install -r requirements.txt

On Colab, skip the PyTorch installation if torch.__version__ is already 2.11.0+cu128. The requirements record the tested versions; package availability depends on platform and Python compatibility. CPU inference is not implemented by this script.

Classify a URL string

python inference.py --revision c1c2501da8ab5bf02d89d3bf2d6f04719faf525a --url "https://example.com/account"

The script prints JSON containing:

  • label: legitimate or phishing.
  • prediction: 1 for legitimate or 0 for phishing.
  • phishing_score_uncalibrated: the relative softmax score for answer B.
  • phishing_logit_margin: the B logit minus the A logit.
  • url_truncated: whether the URL was shortened for the input budget.

The example uses a placeholder URL. Its prediction is not a safety assessment. The script does not visit the URL.

Scores are not calibrated confidence estimates. A legitimate prediction does not establish that a website is safe.

Input checks are basic, not comprehensive URL validation or adversarial sanitization. This script is not a hardened public API.

Model and training

Setting Value
Base checkpoint ibm-granite/granite-4.2-3b
Base revision e459acceac81e5fe67c07d9cfc72329a332e7eb1
Training examples 6,000
Source/class sampling 1,500 examples per source/class combination
Method 4-bit NF4 QLoRA with double quantization
Compute dtype BF16
LoRA targets q_proj, v_proj
LoRA rank / alpha / dropout 8 / 16 / 0.05
Trainable parameters 2,621,440
Epochs 1
Learning rate 0.0001
Warmup 12 optimizer steps
Micro-batch / gradient accumulation 4 / 4
Optimizer steps 375
Sequence limit 512 tokens, including the answer
Hardware NVIDIA A100-SXM4-40GB
Training duration 14.96 minutes
Observed peak allocated training memory 7.03 GiB

Only the answer token contributes to the causal-language-model loss. The quantized base weights remain frozen.

Scope and design choices

Why q_proj and v_proj? Rank-eight adapters on these attention projections provided a small, affordable starting configuration. Only 2,621,440 parameters were trainable. This was a budget choice, not a claim that these are the best target modules; broader adapter coverage was not tested.

Why one epoch? The experiment used one preselected training run to keep cost and iteration manageable. It did not compare epoch counts or establish convergence, and no additional epochs were selected after inspecting final test results.

Why a lightweight baseline? TF-IDF trained on the same 6,000 URLs in seconds. It tests whether the LLM's additional complexity provides an advantage under the same evaluation conditions.

Inputs and scoring

The only predictive input is the URL string. The model does not fetch webpages or inspect HTML, DNS, redirects, reputation feeds, or certificates.

The Granite chat template is used with enable_thinking=False.

  • A = legitimate
  • B = phishing
  • The higher next-token logit determines the prediction.
  • Exact ties select A.
  • Dataset numeric labels are 1 = legitimate and 0 = phishing.

Relative A/B softmax scores are not calibrated phishing probabilities. No classification threshold was tuned against final test results.

The prompt is stored in experiment_config/prompt_config.json. Instructions and the assistant prefix are preserved when long URLs are shortened to fit the token budget. Two training URLs were truncated; one internal-test URL and one published-subset URL were truncated. No external-test URLs were truncated.

Arbitrary generation through a text-generation widget is not the evaluated classification protocol.

Data and evaluation design

Training uses raw URLs from PhiUSIIL and the published training partition of PhreshPhish.

Preparation includes missing-value checks, exact-URL deduplication, conflicting-label exclusions, and public registrable-domain grouping. Public suffixes are used without private-suffix tenant separation.

Development splits were assigned globally by public domain before sampling. Training was capped at five URLs per public domain; validation and final benchmark samples use at most one URL per public domain.

PhreshPhish published-test domains were excluded from development. Previously used Experiment 1 evaluation domains were also excluded.

Benchmarks

  1. Internal mixed-source: 400 URLs from the two training sources, on held-out public domains.
  2. PhreshPhish published subset: 500 balanced examples selected from the authors' published test partition after domain exclusions. This is not the authors' full benchmark or its original class prevalence.
  3. External source: 500 balanced examples from a malicious-URL collection excluded from training. Only benign and phishing classes are retained. Public domains found in the full PhiUSIIL source and extracted PhreshPhish train/test metadata are excluded.

All benchmarks contain equal numbers of legitimate and phishing examples. Their class prevalence and domain-filtered composition are not estimates of real-world browsing traffic.

Results

Recall measures phishing detection; the false-positive rate measures legitimate URLs incorrectly flagged as phishing. Precision, F1, and error counts are available in the full results CSV.

Benchmark Model Accuracy Recall False-positive rate ROC-AUC
Internal Unchanged Granite 71.75% 97.50% 54.00% 0.8586
Internal QLoRA Granite 94.50% 93.50% 4.50% 0.9890
Internal TF-IDF + logistic 91.00% 87.50% 5.50% 0.9784
Published subset Unchanged Granite 72.60% 98.40% 53.20% 0.8904
Published subset QLoRA Granite 86.60% 83.60% 10.40% 0.9450
Published subset TF-IDF + logistic 76.00% 64.40% 12.40% 0.8314
External Unchanged Granite 48.60% 92.00% 94.80% 0.5409
External QLoRA Granite 54.40% 72.80% 64.00% 0.6359
External TF-IDF + logistic 55.20% 63.20% 52.80% 0.6029

Benchmark accuracy and false-positive rates

Full metrics and paired-bootstrap results are included in results/.

Accuracy differences versus TF-IDF

Approximate 95% paired-bootstrap intervals, in percentage points:

Benchmark QLoRA minus TF-IDF Interval
Internal +3.50 0.75 to 6.25
Published subset +10.60 7.20 to 14.00
External -0.80 -5.20 to 3.60

The bootstrap uses 5,000 replicates and preserves class/source strata. Intervals are per-comparison, not multiplicity-adjusted. They do not capture training-seed variability, label noise, or all campaign correlations.

The internal advantage over TF-IDF is modest: 3.5 percentage points, with a 95% interval from 0.75 to 6.25. The interval narrowly excludes zero for these sampled examples; it does not establish that the gain would persist across independent training seeds.

The published-subset advantage is larger at 10.6 points (7.2 to 14.0). The external difference is -0.8 points (-5.2 to 3.6), providing no clear accuracy advantage over TF-IDF. Only one training seed was evaluated.

Query-string shortcut check

No legitimate training examples contained query strings. To check whether gains were limited to flagging query-containing URLs, a post-hoc analysis compared models on test URLs without queries.

  • Internal: all 14 net additional correct predictions over TF-IDF came from URLs without queries. On this subset, QLoRA's accuracy advantage was 3.59 percentage points (95% paired-bootstrap interval: 0.77 to 6.41).
  • Published subset: 50 of 53 net additional correct predictions over TF-IDF came from URLs without queries. The no-query advantage was 10.53 points (7.16 to 13.89).
  • External: no-query accuracy remained poor: 54.27% for QLoRA versus 55.77% for TF-IDF.

Simply flagging query-containing test URLs cannot explain the familiar-source gains. This does not establish causal independence from query-related learning or rule out other shortcuts. Intervals use 5,000 paired, source/class-stratified bootstrap replicates and are not adjusted for multiple comparisons or training-seed variability.

Release verification

The uploaded adapter and inference script were verified at revision c1c2501da8ab5bf02d89d3bf2d6f04719faf525a on an NVIDIA A100-SXM4-40GB with PyTorch 2.11.0+cu128.

All 1,400 benchmark predictions, truncation flags, and label-logit margins matched the recorded fine-tuned evaluations exactly:

  • Internal benchmark: 400 examples.
  • Published PhreshPhish subset: 500 examples.
  • External benchmark: 500 examples.

Inference explicitly uses BF16 autocast. An earlier loader without this context produced four different predictions. The corrected revision reproduced the original results without retraining.

The aggregate report is in results/adapter_verification.json. Exact numerical agreement on other hardware, batch sizes, or dependency versions is not guaranteed.

The training memory figure is an observed measurement, not a certified minimum inference requirement. The installation commands describe the recorded environment; every clean platform installation has not been tested.

Verification and limitations

Frozen file checksums, prediction alignment, label mapping, and saved domain separation were verified. Fresh unchanged-model loading reproduced all 400 validation predictions and logit margins exactly.

Important limitations:

  • Cross-source performance remains poor.
  • Training URLs still contain source-specific formatting patterns.
  • No legitimate sampled training URLs contain query strings.
  • Ground-truth labels were not independently verified by visiting sites.
  • Related campaigns may span different domains.
  • Pretraining exposure to public datasets cannot be ruled out.
  • Only one training seed and one main training configuration were tested.
  • The external source was already explored during Experiment 1, although Experiment 2 uses new eligible domains.
  • Experiment 1 and Experiment 2 change multiple variables and are not a controlled model-version or training-data ablation.
  • Adversarial input robustness and operational deployment are untested.

A later audit found that Experiment 1's nonmissing class counts exactly match the first 499,276 records of the complete external CSV; its reported extra row had a missing label. This strongly suggests a partial-prefix read, potentially during upload, although the precise cause was not recorded. Three independent parsing paths agree that the completed file contains 651,191 records. Experiment 2 used the completed file with reconciled counts and verified split reconstruction.

See the row-count audit.

External CSV SHA256: d83ce942075dd63ed4d11560cfdcd9d512caa3d680e292f22cab484e8f074d01

Data sources and attribution

Raw dataset files are not redistributed in this repository. The linked dataset records provide source documentation and applicable terms. Dataset terms are separate from this adapter's Apache 2.0 license.

Licensing

The adapter and original project code are released under Apache 2.0. The IBM Granite base model is separately available under Apache 2.0. The Apache 2.0 license text is included in LICENSE. Upstream legal files found at the pinned revision are preserved under upstream/; see ATTRIBUTION.md.

Dataset licenses and terms apply separately. Raw datasets are not redistributed here. See the data-source attribution above.

Intended use

Educational research into URL classification, parameter-efficient fine-tuning, dataset shortcuts, and cross-source generalization.

Not intended for automated blocking, declaring URLs safe, or replacing operational anti-phishing defenses.

Public result reproduction

Per-example predictions for all 1,400 final benchmark examples are published in results/predictions/. They include URL and domain hashes, original labels, source strata, all three predictions, Granite logit margins, and saved model scores.

These files reproduce all nine metric rows and all six paired-bootstrap comparisons without private Drive files or raw URLs. Hashes permit membership testing and are not strong anonymization.

From a downloaded copy of the current repository:

python -m pip install -r scripts/requirements-analysis.txt
python scripts/analyze_public.py --predictions results/predictions --output analysis

The older evaluate.py and bootstrap.py retain the original-project audit workflow. Use analyze_public.py for the public prediction tables. Raw URLs are still required to regenerate model predictions. This closes public metric/bootstrap reproduction, not full training or source-label verification.

Inference protocol

The verified benchmark protocol uses batch size 4, left padding, BF16 autocast, and the original row order published in the prediction tables. Batch composition and batch-size invariance were not tested. A single-URL CLI call creates a one-example batch and is a usage example, not an assertion of identical benchmark logits.

The base checkpoint is a dense Granite transformer, not a Mamba hybrid. No position-ID correction has been applied or claimed necessary. A change in margins across batch sizes alone would not distinguish floating-point/kernel differences from a positional implementation issue.

The helper prepare_model_for_kbit_training is retained to reproduce the base-weight casting used during training. This inference call does not perform training and disables gradient-checkpointing setup.

adapter_config.json pins the base revision to e459acceac81e5fe67c07d9cfc72329a332e7eb1. This pin does not force 4-bit NF4 loading, BF16 autocast, the evaluated prompt, or direct A/B scoring. Hugging Face sidebar snippets and the text-generation widget are not the evaluated protocol. Use inference.py and its recorded configuration.

Downloads last month
90
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bengoldberg0/granite-4.2-3b-phishing-url-qlora

Adapter
(4)
this model

Dataset used to train bengoldberg0/granite-4.2-3b-phishing-url-qlora