Instructions to use bengoldberg0/granite-4.2-3b-phishing-url-qlora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bengoldberg0/granite-4.2-3b-phishing-url-qlora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("ibm-granite/granite-4.2-3b") model = PeftModel.from_pretrained(base_model, "bengoldberg0/granite-4.2-3b-phishing-url-qlora") - Notebooks
- Google Colab
- Kaggle
- Granite 4.2 3B - Phishing URL Classification with QLoRA
Granite 4.2 3B - Phishing URL Classification with QLoRA
94.5% accuracy on held-out domains from the training sources; 54.4% on a separate external source. This study examines why strong benchmark performance did not translate into reliable cross-source detection.
Research use only. The external false-positive rate was 64.0%. Scores are uncalibrated and should not be used to decide whether a URL is safe.
Training and evaluation methods | Results | Inference code
Project summary
This project tests whether QLoRA improves URL-only phishing classification and whether the gains carry across data sources.
Experiment 1 fine-tuned Granite 4.0 on PhiUSIIL alone. It reached 99.6% internal accuracy but labeled every external test example as phishing. Audits exposed strong differences in URL structure between the sources.
Experiment 2, released here, fine-tuned Granite 4.2 on 6,000 URLs from PhiUSIIL and PhreshPhish. It compared unchanged Granite, QLoRA, and TF-IDF across internal, published-partition, and external-source tests. The model, training data, and test selections changed between experiments, so this is not a controlled data-only or model-only ablation.
The work centers on data preparation, public-domain exclusions, paired evaluation, and reproduction of the released artifact. The adapter is one output of that experiment.
This is an independent student project, not an IBM-endorsed detector.
Pipeline
The experiment code is public:
- Data preparation
- Metadata retrieval
- QLoRA training
- TF-IDF training
- Metric evaluation
- Paired bootstrap
See PIPELINE.md for commands, dependencies, and verification scope. Preparation, baseline training, and analysis were checked against the original artifacts. The refactored Granite training command was not rerun.
Public per-example prediction tables and scripts/analyze_public.py
reproduce the reported metrics and paired bootstrap without private files.
Regenerating predictions and training still requires the source data.
Usage
This repository contains a PEFT adapter in Safetensors format, not a standalone model. The script downloads the pinned IBM base checkpoint and applies the adapter.
The verified environment used Linux, Python 3.13, PyTorch 2.11.0+cu128, and an NVIDIA A100 40 GB. The script requires a CUDA GPU with BF16 support. Identical outputs on other hardware or software versions are not guaranteed.
Installation
Use an isolated environment with a compatible NVIDIA driver:
python -m pip install "torch==2.11.0" --index-url https://download.pytorch.org/whl/cu128
python -m pip install "huggingface_hub==1.31.0"
hf download bengoldberg0/granite-4.2-3b-phishing-url-qlora inference.py requirements.txt --revision c1c2501da8ab5bf02d89d3bf2d6f04719faf525a --local-dir phishing-url-model
cd phishing-url-model
python -m pip install -r requirements.txt
On Colab, skip the PyTorch installation if torch.__version__ is already
2.11.0+cu128. The requirements record the tested versions; package
availability depends on platform and Python compatibility.
CPU inference is not implemented by this script.
Classify a URL string
python inference.py --revision c1c2501da8ab5bf02d89d3bf2d6f04719faf525a --url "https://example.com/account"
The script prints JSON containing:
label:legitimateorphishing.prediction:1for legitimate or0for phishing.phishing_score_uncalibrated: the relative softmax score for answer B.phishing_logit_margin: the B logit minus the A logit.url_truncated: whether the URL was shortened for the input budget.
The example uses a placeholder URL. Its prediction is not a safety assessment. The script does not visit the URL.
Scores are not calibrated confidence estimates. A legitimate prediction does not establish that a website is safe.
Input checks are basic, not comprehensive URL validation or adversarial sanitization. This script is not a hardened public API.
Model and training
| Setting | Value |
|---|---|
| Base checkpoint | ibm-granite/granite-4.2-3b |
| Base revision | e459acceac81e5fe67c07d9cfc72329a332e7eb1 |
| Training examples | 6,000 |
| Source/class sampling | 1,500 examples per source/class combination |
| Method | 4-bit NF4 QLoRA with double quantization |
| Compute dtype | BF16 |
| LoRA targets | q_proj, v_proj |
| LoRA rank / alpha / dropout | 8 / 16 / 0.05 |
| Trainable parameters | 2,621,440 |
| Epochs | 1 |
| Learning rate | 0.0001 |
| Warmup | 12 optimizer steps |
| Micro-batch / gradient accumulation | 4 / 4 |
| Optimizer steps | 375 |
| Sequence limit | 512 tokens, including the answer |
| Hardware | NVIDIA A100-SXM4-40GB |
| Training duration | 14.96 minutes |
| Observed peak allocated training memory | 7.03 GiB |
Only the answer token contributes to the causal-language-model loss. The quantized base weights remain frozen.
Scope and design choices
Why q_proj and v_proj? Rank-eight adapters on these attention projections provided a small, affordable starting configuration. Only 2,621,440 parameters were trainable. This was a budget choice, not a claim that these are the best target modules; broader adapter coverage was not tested.
Why one epoch? The experiment used one preselected training run to keep cost and iteration manageable. It did not compare epoch counts or establish convergence, and no additional epochs were selected after inspecting final test results.
Why a lightweight baseline? TF-IDF trained on the same 6,000 URLs in seconds. It tests whether the LLM's additional complexity provides an advantage under the same evaluation conditions.
Inputs and scoring
The only predictive input is the URL string. The model does not fetch webpages or inspect HTML, DNS, redirects, reputation feeds, or certificates.
The Granite chat template is used with enable_thinking=False.
A= legitimateB= phishing- The higher next-token logit determines the prediction.
- Exact ties select
A. - Dataset numeric labels are
1= legitimate and0= phishing.
Relative A/B softmax scores are not calibrated phishing probabilities. No classification threshold was tuned against final test results.
The prompt is stored in experiment_config/prompt_config.json.
Instructions and the assistant prefix are preserved when long URLs are
shortened to fit the token budget. Two training URLs were truncated;
one internal-test URL and one published-subset URL were truncated.
No external-test URLs were truncated.
Arbitrary generation through a text-generation widget is not the evaluated classification protocol.
Data and evaluation design
Training uses raw URLs from PhiUSIIL and the published training partition of PhreshPhish.
Preparation includes missing-value checks, exact-URL deduplication, conflicting-label exclusions, and public registrable-domain grouping. Public suffixes are used without private-suffix tenant separation.
Development splits were assigned globally by public domain before sampling. Training was capped at five URLs per public domain; validation and final benchmark samples use at most one URL per public domain.
PhreshPhish published-test domains were excluded from development. Previously used Experiment 1 evaluation domains were also excluded.
Benchmarks
- Internal mixed-source: 400 URLs from the two training sources, on held-out public domains.
- PhreshPhish published subset: 500 balanced examples selected from the authors' published test partition after domain exclusions. This is not the authors' full benchmark or its original class prevalence.
- External source: 500 balanced examples from a malicious-URL collection excluded from training. Only benign and phishing classes are retained. Public domains found in the full PhiUSIIL source and extracted PhreshPhish train/test metadata are excluded.
All benchmarks contain equal numbers of legitimate and phishing examples. Their class prevalence and domain-filtered composition are not estimates of real-world browsing traffic.
Results
Recall measures phishing detection; the false-positive rate measures legitimate URLs incorrectly flagged as phishing. Precision, F1, and error counts are available in the full results CSV.
| Benchmark | Model | Accuracy | Recall | False-positive rate | ROC-AUC |
|---|---|---|---|---|---|
| Internal | Unchanged Granite | 71.75% | 97.50% | 54.00% | 0.8586 |
| Internal | QLoRA Granite | 94.50% | 93.50% | 4.50% | 0.9890 |
| Internal | TF-IDF + logistic | 91.00% | 87.50% | 5.50% | 0.9784 |
| Published subset | Unchanged Granite | 72.60% | 98.40% | 53.20% | 0.8904 |
| Published subset | QLoRA Granite | 86.60% | 83.60% | 10.40% | 0.9450 |
| Published subset | TF-IDF + logistic | 76.00% | 64.40% | 12.40% | 0.8314 |
| External | Unchanged Granite | 48.60% | 92.00% | 94.80% | 0.5409 |
| External | QLoRA Granite | 54.40% | 72.80% | 64.00% | 0.6359 |
| External | TF-IDF + logistic | 55.20% | 63.20% | 52.80% | 0.6029 |
Full metrics and paired-bootstrap results are included in results/.
Accuracy differences versus TF-IDF
Approximate 95% paired-bootstrap intervals, in percentage points:
| Benchmark | QLoRA minus TF-IDF | Interval |
|---|---|---|
| Internal | +3.50 | 0.75 to 6.25 |
| Published subset | +10.60 | 7.20 to 14.00 |
| External | -0.80 | -5.20 to 3.60 |
The bootstrap uses 5,000 replicates and preserves class/source strata. Intervals are per-comparison, not multiplicity-adjusted. They do not capture training-seed variability, label noise, or all campaign correlations.
The internal advantage over TF-IDF is modest: 3.5 percentage points, with a 95% interval from 0.75 to 6.25. The interval narrowly excludes zero for these sampled examples; it does not establish that the gain would persist across independent training seeds.
The published-subset advantage is larger at 10.6 points (7.2 to 14.0). The external difference is -0.8 points (-5.2 to 3.6), providing no clear accuracy advantage over TF-IDF. Only one training seed was evaluated.
Query-string shortcut check
No legitimate training examples contained query strings. To check whether gains were limited to flagging query-containing URLs, a post-hoc analysis compared models on test URLs without queries.
- Internal: all 14 net additional correct predictions over TF-IDF came from URLs without queries. On this subset, QLoRA's accuracy advantage was 3.59 percentage points (95% paired-bootstrap interval: 0.77 to 6.41).
- Published subset: 50 of 53 net additional correct predictions over TF-IDF came from URLs without queries. The no-query advantage was 10.53 points (7.16 to 13.89).
- External: no-query accuracy remained poor: 54.27% for QLoRA versus 55.77% for TF-IDF.
Simply flagging query-containing test URLs cannot explain the familiar-source gains. This does not establish causal independence from query-related learning or rule out other shortcuts. Intervals use 5,000 paired, source/class-stratified bootstrap replicates and are not adjusted for multiple comparisons or training-seed variability.
Release verification
The uploaded adapter and inference script were verified at revision
c1c2501da8ab5bf02d89d3bf2d6f04719faf525a on an NVIDIA A100-SXM4-40GB
with PyTorch 2.11.0+cu128.
All 1,400 benchmark predictions, truncation flags, and label-logit margins matched the recorded fine-tuned evaluations exactly:
- Internal benchmark: 400 examples.
- Published PhreshPhish subset: 500 examples.
- External benchmark: 500 examples.
Inference explicitly uses BF16 autocast. An earlier loader without this context produced four different predictions. The corrected revision reproduced the original results without retraining.
The aggregate report is in results/adapter_verification.json.
Exact numerical agreement on other hardware, batch sizes, or dependency
versions is not guaranteed.
The training memory figure is an observed measurement, not a certified minimum inference requirement. The installation commands describe the recorded environment; every clean platform installation has not been tested.
Verification and limitations
Frozen file checksums, prediction alignment, label mapping, and saved domain separation were verified. Fresh unchanged-model loading reproduced all 400 validation predictions and logit margins exactly.
Important limitations:
- Cross-source performance remains poor.
- Training URLs still contain source-specific formatting patterns.
- No legitimate sampled training URLs contain query strings.
- Ground-truth labels were not independently verified by visiting sites.
- Related campaigns may span different domains.
- Pretraining exposure to public datasets cannot be ruled out.
- Only one training seed and one main training configuration were tested.
- The external source was already explored during Experiment 1, although Experiment 2 uses new eligible domains.
- Experiment 1 and Experiment 2 change multiple variables and are not a controlled model-version or training-data ablation.
- Adversarial input robustness and operational deployment are untested.
A later audit found that Experiment 1's nonmissing class counts exactly match the first 499,276 records of the complete external CSV; its reported extra row had a missing label. This strongly suggests a partial-prefix read, potentially during upload, although the precise cause was not recorded. Three independent parsing paths agree that the completed file contains 651,191 records. Experiment 2 used the completed file with reconciled counts and verified split reconstruction.
See the row-count audit.
External CSV SHA256:
d83ce942075dd63ed4d11560cfdcd9d512caa3d680e292f22cab484e8f074d01
Data sources and attribution
- IBM Granite 4.2: https://huggingface.co/ibm-granite/granite-4.2-3b
- PhiUSIIL, original dataset record: https://archive.ics.uci.edu/dataset/967/phiusiil+phishing+url+dataset
- PhreshPhish: https://huggingface.co/datasets/phreshphish/phreshphish
- PhreshPhish revision:
eabec4b7a66324b79cc8a0ad856d1731dc26fe1a - External malicious-URL dataset: https://www.kaggle.com/datasets/sid321axn/malicious-urls-dataset
Raw dataset files are not redistributed in this repository. The linked dataset records provide source documentation and applicable terms. Dataset terms are separate from this adapter's Apache 2.0 license.
Licensing
The adapter and original project code are released under Apache 2.0.
The IBM Granite base model is separately available under Apache 2.0.
The Apache 2.0 license text is included in LICENSE. Upstream legal files found at the pinned revision are preserved under upstream/; see ATTRIBUTION.md.
Dataset licenses and terms apply separately. Raw datasets are not redistributed here. See the data-source attribution above.
Intended use
Educational research into URL classification, parameter-efficient fine-tuning, dataset shortcuts, and cross-source generalization.
Not intended for automated blocking, declaring URLs safe, or replacing operational anti-phishing defenses.
Public result reproduction
Per-example predictions for all 1,400 final benchmark examples are published in results/predictions/. They include URL and domain hashes, original labels, source strata, all three predictions, Granite logit margins, and saved model scores.
These files reproduce all nine metric rows and all six paired-bootstrap comparisons without private Drive files or raw URLs. Hashes permit membership testing and are not strong anonymization.
From a downloaded copy of the current repository:
python -m pip install -r scripts/requirements-analysis.txt
python scripts/analyze_public.py --predictions results/predictions --output analysis
The older evaluate.py and bootstrap.py retain the original-project
audit workflow. Use analyze_public.py for the public prediction tables.
Raw URLs are still required to regenerate model predictions.
This closes public metric/bootstrap reproduction, not full training
or source-label verification.
Inference protocol
The verified benchmark protocol uses batch size 4, left padding, BF16 autocast, and the original row order published in the prediction tables. Batch composition and batch-size invariance were not tested. A single-URL CLI call creates a one-example batch and is a usage example, not an assertion of identical benchmark logits.
The base checkpoint is a dense Granite transformer, not a Mamba hybrid. No position-ID correction has been applied or claimed necessary. A change in margins across batch sizes alone would not distinguish floating-point/kernel differences from a positional implementation issue.
The helper prepare_model_for_kbit_training is retained to reproduce
the base-weight casting used during training. This inference call
does not perform training and disables gradient-checkpointing setup.
adapter_config.json pins the base revision to e459acceac81e5fe67c07d9cfc72329a332e7eb1.
This pin does not force 4-bit NF4 loading, BF16 autocast, the evaluated
prompt, or direct A/B scoring. Hugging Face sidebar snippets and the
text-generation widget are not the evaluated protocol. Use inference.py
and its recorded configuration.
- Downloads last month
- 90
Model tree for bengoldberg0/granite-4.2-3b-phishing-url-qlora
Base model
ibm-granite/granite-4.1-3b-base
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("ibm-granite/granite-4.2-3b") model = PeftModel.from_pretrained(base_model, "bengoldberg0/granite-4.2-3b-phishing-url-qlora")