# Training and evaluation methods This page documents Experiment 2. Refactored preparation, training, evaluation, and bootstrap code is available in [scripts/](scripts/). See [PIPELINE.md](PIPELINE.md) for commands and the verification scope. Original development notebooks are not distributed. Public URL-hashed prediction tables support metric and bootstrap reproduction; regenerating predictions and training still requires source data. ## Data preparation Training candidates came from PhiUSIIL and the published training partition of PhreshPhish. Only URL text was used as predictive input. HTML, source identifiers, dates, and domain groups were not model inputs. Preparation steps: 1. Remove missing or empty URLs and strip surrounding whitespace. 2. Map labels to 0 for phishing and 1 for legitimate. 3. Exclude every occurrence of an exact URL with conflicting labels. 4. Deduplicate exact URLs across sources and retain source membership. 5. Parse public registrable domains with tldextract, with private suffixes disabled. Exclude URLs that cannot be parsed. 6. Exclude domains used in Experiment 1 validation and test sets. 7. Exclude PhreshPhish published-test domains from development candidates. Exact duplicates shared between sources retained the PhreshPhish record and its collection date, while preserving both source memberships. Paths, query strings, and schemes were not removed or rewritten. ## Fixed development splits Domains were assigned globally, across sources, using SHA256 of `42:domain`. The first eight digest bytes were interpreted as a big-endian integer modulo 100. Buckets below 80 went to training, 80 through 89 to validation, and 90 through 99 to internal testing. Rows were shuffled before applying a global per-domain cap. Training used at most five URLs per public domain. Validation and internal testing used at most one. | Split | URLs | Examples per source/class combination | Sampling seed | |---|---:|---:|---:| | Training | 6,000 | 1,500 | 42 | | Validation | 400 | 100 | 43 | | Internal test | 400 | 100 | 44 | Splits were saved with SHA256 checksums. URL and public-domain separation were checked before evaluation. ## Models Granite 4.2 used a fixed non-thinking chat prompt and two answer labels: `A` for legitimate and `B` for phishing. Classification compared their next-token logits, with ties selecting `A`. QLoRA used 4-bit NF4 weights with double quantization, BF16 computation, rank 8 adapters on q_proj and v_proj, alpha 16, and dropout 0.05. Only the final answer token contributed to the training loss. One epoch used 6,000 examples, a micro-batch of four, gradient accumulation of four, learning rate 0.0001, and 12 warmup steps. The input limit was 512 tokens including the answer. Long inputs retained a prefix of the URL while preserving instructions and the assistant-answer position. Before/after scoring used the same policy. TF-IDF used character n-grams of lengths 3 to 5, min_df=2, a maximum of 100,000 features, and no automatic lowercasing. Logistic regression used C=1.0, the liblinear solver, max_iter=1000, and random_state=42. It was trained on the same 6,000 URLs as Granite. No classification thresholds were tuned against final tests. ## Final benchmarks - Internal: 400 examples from the familiar training sources, with public domains held out from training and validation. - Published subset: 250 examples per class from PhreshPhish's published test partition, after cleaning and previous-evaluation exclusions. One URL per public domain was selected using seed 2027. - External: 250 examples per class from the malicious-URL source, using seed 2028 and one URL per public domain. Malware and defacement classes were excluded after exact-URL conflict checks. External selection excluded public domains appearing in the full PhiUSIIL file, PhreshPhish train/test metadata, and Experiment 1's external test. Exact-string exclusion also covered reference URLs whose domains could not be parsed. All benchmarks are balanced, filtered samples, not estimates of real-world phishing prevalence. The external source had already been examined in Experiment 1, although these selected domains were new relative to its evaluation sets. ## Metrics and uncertainty Metrics were recomputed from saved predictions after checking row alignment and label mappings. Reports include accuracy, phishing precision, recall, F1, false-positive rate, and ROC-AUC. Granite ROC-AUC used the B-minus-A logit margin rather than rounded softmax scores. The classifier's scores are not calibrated probabilities. Accuracy differences used 5,000 paired-bootstrap replicates with seed 42. Resampling preserved source/class strata. Both models used the same sampled indices in each replicate. Intervals are per-comparison and are not multiplicity-adjusted. They do not measure training-seed variability, label uncertainty, or all correlations between campaigns on different domains. ## Reproduction Fresh baseline loading reproduced the recorded validation outputs. The released adapter and inference script reproduced all 1,400 final predictions, truncation flags, and logit margins exactly on the recorded A100 environment after matching BF16 autocast. See [the verification report](results/adapter_verification.json), [metric table](results/verified_metrics.csv), and [paired comparisons](results/paired_accuracy_bootstrap.csv). ## Limits The structural mix improved over Experiment 1, but source-specific patterns remained. No legitimate sampled training URLs contained queries, and most used HTTPS. Strong familiar-source results did not transfer to the external source. The saved artifacts establish reproducible inference on the tested environment. They do not independently verify dataset labels, prove the absence of pretraining exposure, or establish deployment readiness.