bengoldberg0 commited on
Commit
cc7e628
·
verified ·
1 Parent(s): d35e592

Add query-string subset diagnostic

Browse files
Files changed (1) hide show
  1. README.md +22 -0
README.md CHANGED
@@ -262,6 +262,28 @@ The published-subset advantage is larger at 10.6 points
262
  (-5.2 to 3.6), providing no clear accuracy advantage over TF-IDF.
263
  Only one training seed was evaluated.
264
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
265
  ## Release verification
266
 
267
  The uploaded adapter and inference script were verified at revision
 
262
  (-5.2 to 3.6), providing no clear accuracy advantage over TF-IDF.
263
  Only one training seed was evaluated.
264
 
265
+ ### Query-string shortcut check
266
+
267
+ No legitimate training examples contained query strings. To check
268
+ whether gains were limited to flagging query-containing URLs, a
269
+ post-hoc analysis compared models on test URLs without queries.
270
+
271
+ - **Internal:** all 14 net additional correct predictions over TF-IDF
272
+ came from URLs without queries. On this subset, QLoRA's accuracy
273
+ advantage was 3.59 percentage points (95% paired-bootstrap interval:
274
+ 0.77 to 6.41).
275
+ - **Published subset:** 50 of 53 net additional correct predictions
276
+ over TF-IDF came from URLs without queries. The no-query advantage
277
+ was 10.53 points (7.16 to 13.89).
278
+ - **External:** no-query accuracy remained poor: 54.27% for QLoRA
279
+ versus 55.77% for TF-IDF.
280
+
281
+ Simply flagging query-containing test URLs cannot explain the
282
+ familiar-source gains. This does not establish causal independence
283
+ from query-related learning or rule out other shortcuts. Intervals
284
+ use 5,000 paired, source/class-stratified bootstrap replicates and
285
+ are not adjusted for multiple comparisons or training-seed variability.
286
+
287
  ## Release verification
288
 
289
  The uploaded adapter and inference script were verified at revision