TareHimself commited on
Commit
582850a
·
verified ·
1 Parent(s): f2241f6

from 20260830-030109_syn-real-v2

Browse files
Files changed (5) hide show
  1. README.md +61 -0
  2. config.json +20 -0
  3. model.pt +3 -0
  4. model.safetensors +3 -0
  5. tm_meta.json +22 -0
README.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: segmentation-models-pytorch
4
+ pipeline_tag: image-segmentation
5
+ tags:
6
+ - comics
7
+ - manga
8
+ - text-segmentation
9
+ language:
10
+ - en
11
+ - ja
12
+ - ko
13
+ - zh
14
+ ---
15
+
16
+ # comic-text-mask
17
+
18
+ Binary **text-mask segmenter** for the
19
+ [comic-localizer](https://github.com/TareHimself/comic-localizer) cleaning
20
+ pipeline. It runs on a detector's text-region crop and returns a per-pixel
21
+ "text vs not-text" mask that feeds LaMa inpainting. Locating and grouping text
22
+ is the detector's job, not this model's.
23
+
24
+ Training data (see the [training repo](https://github.com/TareHimself/comic-localizer-text-masking)): synthetic text rendered onto
25
+ (a) procedurally generated flat surfaces with synthetic clutter and (b) real
26
+ cleaned comic pages, framed as detector-style crops. Text is Latin, Japanese
27
+ (kana + kanji, horizontal and vertical), Korean, and Chinese, with a large
28
+ fraction of random-glyph runs so rare characters are covered. No source imagery
29
+ is redistributed.
30
+
31
+ Validation (held-out synthetic + real crops): IoU 0.918,
32
+ precision 0.956, recall 0.959.
33
+
34
+ ## Use
35
+
36
+ ```python
37
+ import numpy as np, torch
38
+ from huggingface_hub import hf_hub_download
39
+
40
+ meta = json.load(open(hf_hub_download("TareHimself/comic-text-mask", "tm_meta.json")))
41
+ model = torch.jit.load(hf_hub_download("TareHimself/comic-text-mask", "model.pt")).eval()
42
+ S = meta["imgsz"]
43
+
44
+ # letterbox `rgb` (H,W,3 uint8) into an SxS square, pad 0, keep the paste box
45
+ # ... then:
46
+ x = torch.from_numpy(square).permute(2, 0, 1).unsqueeze(0) # (1,3,S,S) uint8
47
+ prob = model(x)[0, 0].numpy() # (S,S) float
48
+ mask = (prob > meta["threshold"]).astype("uint8") * 255
49
+ # crop the paste box back out and resize to the original size
50
+ ```
51
+
52
+ Or load the raw weights with `segmentation-models-pytorch`:
53
+
54
+ ```python
55
+ import segmentation_models_pytorch as smp
56
+ model = smp.from_pretrained("TareHimself/comic-text-mask") # normalisation NOT baked in
57
+ ```
58
+
59
+ `model.pt` has `/255`, ImageNet normalisation, and the final sigmoid baked into
60
+ the graph; it expects letterboxed **uint8** RGB. `model.safetensors` is the
61
+ pristine network.
config.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_model_class": "Unet",
3
+ "activation": null,
4
+ "aux_params": null,
5
+ "classes": 1,
6
+ "decoder_attention_type": null,
7
+ "decoder_channels": [
8
+ 256,
9
+ 128,
10
+ 64,
11
+ 32,
12
+ 16
13
+ ],
14
+ "decoder_interpolation": "nearest",
15
+ "decoder_use_norm": "batchnorm",
16
+ "encoder_depth": 5,
17
+ "encoder_name": "resnet18",
18
+ "encoder_weights": null,
19
+ "in_channels": 3
20
+ }
model.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:37fa988fcf2908c23d75af3988b00fd3210027627052b12a5abf606ac2d65da7
3
+ size 57586952
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0d1299ce73b89cf7065418f00b3d4806b8f73910299fd30c41e83a3463c4e185
3
+ size 57377484
tm_meta.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "imgsz": 384,
3
+ "input": "letterboxed uint8 RGB (1,3,imgsz,imgsz), pad 0",
4
+ "output": "float32 text probability (1,1,imgsz,imgsz)",
5
+ "normalization": "baked into model.pt",
6
+ "threshold": 0.5,
7
+ "arch": "unet",
8
+ "encoder": "resnet18",
9
+ "metrics": {
10
+ "iou": 0.91798,
11
+ "precision": 0.95562,
12
+ "recall": 0.95885
13
+ },
14
+ "languages": [
15
+ "en_GB",
16
+ "en_US",
17
+ "ja",
18
+ "ko",
19
+ "zh_CN",
20
+ "zh_TW"
21
+ ]
22
+ }