# Experiment pipeline These scripts were extracted or refactored from the executed Experiment 2 notebook. They expose the preparation, training, evaluation, and bootstrap implementation. ## Verification scope | Component | Checked result | |---|---| | Metadata download | 666,315 rows matched the original cache | | Preparation | All five ordered splits matched | | TF-IDF training | All 1,800 validation/test predictions matched | | Evaluation | All nine metric rows matched | | Bootstrap | All six paired comparisons matched | | Granite preparation | All 6,000 training encodings and collator checks matched | | Refactored Granite training | Configuration checked; training was not rerun | Reports are in `pipeline_verification/`. The released adapter's separate inference verification is documented in the model card. It is not proof that the refactored training command will produce bitwise-identical adapter weights. ## Obtain the repository Download or clone the current repository, including `scripts/` and `pipeline_config/`. Run the commands below from the repository root. Record the repository commit used for your run. Do not use the older inference-verification commit to download these pipeline scripts; they were added afterward. ## Source inputs Obtain the original PhiUSIIL URL/label CSV and malicious-URL CSV from the sources linked in the model card, subject to their terms. The expected filenames and byte hashes are in `pipeline_config/data_sources.json`. The recorded PhiUSIIL input is the two-column CSV used in the experiment. An independently exported equivalent CSV can have a different byte hash. The script deliberately rejects mismatched input hashes; a different export requires an explicit provenance update and fresh validation. PhreshPhish metadata is downloaded from its pinned revision. HTML is not selected. Prior evaluation exclusions are stored as SHA256 hashes of canonical public domains. These hashes are not strong anonymization. ## Download metadata and prepare splits ```bash python -m pip install -r scripts/requirements-data.txt python scripts/download_metadata.py --config pipeline_config/data_sources.json --output work/phreshphish python scripts/prepare_data.py --config pipeline_config --phiusiil inputs/PhiUSIIL_Phishing_URL_Dataset_only_url_label.csv --external inputs/malicious_phish.csv --phresh-root work/phreshphish/eabec4b7a66324b79cc8a0ad856d1731dc26fe1a --output work/reconstructed/data/frozen_v1 ``` Use a new output directory. Preparation checks ordered URL, label, source, domain, and answer fingerprints against the original splits. Parquet bytes may differ across serialization environments. ## Train the lightweight baseline ```bash python -m pip install -r scripts/requirements-baseline.txt python scripts/train_tfidf.py --train work/reconstructed/data/frozen_v1/train.parquet --expected-splits pipeline_config/expected_splits.json --output work/tfidf_run ``` Only load Joblib artifacts from sources you trust. ## Prepare or train Granite Install the recorded CUDA PyTorch build as described in the model card before installing the remaining training dependencies. ```bash python -m pip install -r scripts/requirements-training.txt python scripts/train.py --train work/reconstructed/data/frozen_v1/train.parquet --config pipeline_config --output work/granite_preparation --prepare-only ``` The prepare-only command does not load model weights or train. To launch a new GPU training run, omit `--prepare-only` and use a different output directory: ```bash python scripts/train.py --train work/reconstructed/data/frozen_v1/train.parquet --config pipeline_config --output work/granite_run ``` The training script requires a BF16-capable CUDA GPU. It refuses to overwrite an existing output directory and does not implement automatic resume. ## Recompute the original analysis Per-example predictions for all 1,400 final benchmark examples are published in [results/predictions/](results/predictions/). They include URL and domain hashes, original labels, source strata, all three predictions, Granite logit margins, and saved model scores. These files reproduce all nine metric rows and all six paired-bootstrap comparisons without private Drive files or raw URLs. Hashes permit membership testing and are not strong anonymization. From a downloaded copy of the current repository: ```bash python -m pip install -r scripts/requirements-analysis.txt python scripts/analyze_public.py --predictions results/predictions --output analysis ``` The older `evaluate.py` and `bootstrap.py` retain the original-project audit workflow. Use `analyze_public.py` for the public prediction tables. Raw URLs are still required to regenerate model predictions. This closes public metric/bootstrap reproduction, not full training or source-label verification. ## Files - `scripts/download_metadata.py`: pinned column-selective metadata retrieval. - `scripts/prepare_data.py`: cleaning, exclusions, domain splits, sampling. - `scripts/train.py`: answer-only QLoRA preparation and training. - `scripts/train_tfidf.py`: character TF-IDF and logistic regression. - `scripts/evaluate.py`: saved-prediction checks and metric calculation. - `scripts/bootstrap.py`: paired source/class-stratified bootstrap.