š¾ New (very small) dataset: Cat Keyboard Corpus
My desk is where the sun is, and the cats have right of way. The laptop on it runs a long-term sensor project with an AI assistant, so every so often a cat walks across the keyboard and a message gets sent. Instead of shooing the cats off, I started keeping the messages.
So far: 10 messages, 748 cat keystrokes, 16 distinct keys plus the spacebar, 6 days.
- Favourite key: N (132 presses), then D (108) and B (52). Both leaders arrived today, in two messages nobody saw being typed: a single b followed by 114 n's, then 248 spaces, c, 107 d's, 46 g's, v, b - That second one is a slide across the middle of the keyboard, and the longest message so far (404 characters) - Giddings drifts: his early messages sit on the boer ones on the bottom right next to Enter (k l ; ' , . /) - One message has capital B's: a paw on Shift - One is a collaboration: I typed " mxc" on the wrong keyboard, then Giddings finished it with "ccccccccccccccccccdm"
Each message records who pressed Enter (sent_by: the cat, me, or unknown). I pressed it five times and four are unknown. And today Chester, who walks the keyboard often but had nevere x and pressed Enter himself: the first message acat sent on its own. He was also the first cat the project's camera ever recognized.
Open question for the cat people here: does your cat have a favourite key? Here it's N, D and B, the keys right where a paw lands.
I pulled a dead 2013 Butterfly Labs "Jalapeno" SHA-256 mining ASIC out of a drawer and built a modern Python toolkit for it. Then I hit the wall every honest hardware project should.
I "found" undocumented serial commands the mining software (cgminer) defines but never sends, including a persistent NVRAM scratchpad that survives power cycles. I wrote my name and the repo URL into the silicon; it's still there.
Then I checked prior art. All of it is in Butterfly Labs' own 2012 protocol spec and their open source firmware. Rederivation, not discovery. I even almost filed a "bug" against cgminer before a last look at its code showed it was right and I'd misread it. Twice, the "gotcha" was me not reading carefully.
Why it was still worth it, the firmware source can't tell you the chip still works. So I measured it, model free:
1. Four hours of continuous work, zero compute errors, fully deterministic. 2. The winning nonce count is Poisson(~1), the chip scans the whole 2^32 nonce space per job. 3. Thermally over built: it won't error even with the fan off (~41C max on a desk).
The one genuinely new thing: a dead-core detector. It flags a dead engine as a cold band in the nonce histogram. It can't map the healthy engine partitions (they sum to uniform), only localize the dead ones.
The honest move, go check whether it's already known, costs you a discovery and gives you the truth. Better trade every time.
A dead 2013 Butterfly Labs "Jalapeno" SHA-256 mining ASIC sat in a drawer for a decade. It became the excuse for a small, careful question: how much structure can a tiny, cheap model learn in SHA-256, and how would I know if I were fooling myself? (The ML runs on CPU and a HF job, not the ASIC; the dead miner is just the origin story.)
Three findings, written up honestly:
1. A sharp round-4 cliff. Round-reduced SHA-256 is ~100% distinguishable through 3 rounds, then collapses to chance at round 4 and stays there out to the full 64. Reproduced across 5 seeds.
2. A controls-gated bounded null on full SHA-256: no learnable structure above a ~0.22% resolution floor at n=4,000,000. That is a bounded null at this budget, not a claim that SHA-256 is random.
3. A "signal" in the iterated-hash dynamics that a permuted-label control unmasked as a label-prior artifact. The instrument caught its own false positive. That was the point of building the controls.
Negative results, stated with their resolution. The dataset carries the controls on every row.