Remove training source and make inference self-contained
Remove the bundled training project under source/ while keeping the released model usable through the documented inference example.
inference.py previously downloaded and imported autojev from source/src; simply deleting source/ would break prediction. It now contains the inference-only schema, preprocessing, checkpoint loader, and calibrated readout. Training initialization, checkpoint saving, dataset generation, evaluation, server, configs, and lockfile are removed. Preserve the extracted code's MIT license in LICENSE.inference and update NOTICE.
Model weights, tokenizer/processor files, decision configuration, and release manifest are unchanged. Choice, yes/no, score, and image inference retain their existing interfaces and calculations.
Validation:
- Python 3.12 syntax and git diff whitespace checks pass.
- AST comparison confirms that prompt construction, image handling, answer calculations, and prepare/forward/predict methods match the original release.
- Verified the published PR removes all 24
source/files and preserves every model artifact's Git blob. - Loaded the exact PR revision
d66adf1fae971ad04e04d5f140d64705399b7c2efrom Hugging Face on an NVIDIA H200 viahyperpod_uw1_bo, using the inference script's pinned dependencies (PyTorch2.14.0+cu130, Transformers5.17.0). Slurm job130594completed with exit code 0. - The downloaded checkpoint contains no
source/directory; noautojevmodule was imported. Model parameters have gradients disabled and the model is in evaluation mode. - Validated finite, normalized probabilities and the expected outcomes for all four prediction modes:
| Check | Result |
|---|---|
| Urgent integration support request, yes/no | Urgency probability 0.9930 |
| Route integration failure | technical_support, probability 0.9916 |
| Score a very satisfied customer | 1.9935 on a 0β2 scale |
| Classify a solid red image | red, probability 0.9624 |
Peak allocated GPU memory was 48.71 GiB. Loading took 68.5 seconds; the four inference calls plus validation took 3.5 seconds.
HyperPod environment note: its inherited LD_LIBRARY_PATH prioritizes CUDA 12.8/cuDNN libraries, which conflict with the pinned PyTorch CUDA 13 wheel during image Conv3d. A standalone PyTorch Conv3d reproduced the failure without model code. Clearing LD_LIBRARY_PATH for the test process fixed both that reproducer and the full model test; the PR does not change model code or dependencies to work around the host environment.