vermilion10

Back

vermilion10 · Personal project write-up · September 2026

vermilion10/comic-translator Loading repository info… ★ –⑂ –

Abstract#

Automatic comic translation chains five stages: text detection, recognition, erasure, translation and typesetting. Each stage acts on every region the detector returns, so a false detection is not a harmless extra box. It erases part of the artwork and prints an invented translation on it. We built a comic translator that runs entirely inside a Chrome extension. It covers Japanese, Chinese and Korean sources, and every model runs client-side. We use it to study how text detectors should be evaluated when false positives carry this cost. Seen in that light, the usual evaluation is misleading. Our first fine-tuned detector tripled sound-effect recall at the default threshold, yet at the pipeline’s tolerated false-positive rate of 0.9% it was worse than the model it was meant to replace. We therefore compare detectors at matched spurious-detection ceilings. We also correct a reference set whose ground truth had been pooled from the proposals of the baseline detector and one text finder, and we resample pages and training seeds before believing a ranking. Across six YOLO fine-tunes we found that fixing the labels and mining hard negatives moved the operating curve, while changing the starting checkpoint did not. We also found that seed-to-seed variance was as large as the difference between recipes. We then trained our own detector, TextSeg: a MobileNetV3 and FPN text-probability map with an ignore-aware DBNet-style loss. Combined with negatives mined against TextSeg itself, it beats the baseline detector at the 0.9% ceiling on all three seeds, with page-bootstrap probabilities of 0.92 to 0.97. Its sound-effect recall roughly triples (0.41 against 0.15). Around the detector we report measured choices for per-script OCR, balloon-aware erasing, contrast-guaranteed caption plates and free-tier cloud translation. We designed the evaluation protocol for comic-translation pipelines that erase what they detect, and we expect it to be useful for similar pipelines, though we have tested it only on this one.

Keywords: comic text detection, manga translation, operating-point evaluation, pooling bias, hard-negative mining, in-browser inference


1. Introduction#

1.1 The problem#

A reader who opens a Japanese manga, a Chinese manhua or a Korean webtoon in a browser sees text that most of them cannot read. Tools that translate these images already exist. zyddnys/manga-image-translator [50], comic-translate [49] and koharu [10] follow the same five-stage pattern: detect the lettering, recognise it, erase it, translate it and typeset the translation back into the art. Nearly all of them run as desktop applications or behind a server, so the reader has to save a page, open another program and come back. Hinami et al. [3] built a complete research system around the same pipeline and showed that translation quality depends on the image context as well as the text.

We wanted the same pipeline with no step between reading and translating. That meant a browser extension that works on the page the reader is already viewing. It had to keep the text on the device by default and send nothing to a server unless the reader chose a cloud translator. Every stage therefore runs client-side: ONNX models on onnxruntime-web, OpenCV.js in a sandboxed page, and a quantised language model on WebGPU through WebLLM [25].

1.2 Why detection has to be judged at an operating point#

The detector decides what the rest of the pipeline touches. Each box it emits is cropped for OCR, erased by inpainting, translated and overwritten with English. A missed sound effect stays untranslated, which the reader barely notices. A false box on a face, a heart or a speed line erases part of the drawing and prints an invented translation on it, and the reader notices that at once. So the two error types cost very different amounts. The extension first shipped with a YOLO26n model fine-tuned for manga bubbles [48]. We call it the baseline throughout, because every candidate in this paper is compared against it; TextSeg (Section 6) later replaced it in the extension. On the original 525-region reference, the baseline produced 3 spurious boxes out of 346 detections (0.9%) at its default threshold. We take that rate as the pipeline’s tolerance.

Detection papers for comics usually report average precision, F-measure or pixel-level scores over a fixed test set [4, 6, 53], and model comparisons are usually made at one threshold. That is the right summary when every error costs the same. It is the wrong one here, and our own work shows why. Our first fine-tune raised sound-effect recall at threshold 0.25 from 0.155 to 0.465. It also raised the spurious rate from 0.9% to 7.8%. At a matched spurious rate of 0.9% the fine-tune lost to the baseline on every slice. Most of the gain had come from emitting 462 boxes instead of 346.

Two more problems sit under any such comparison. First, a reference set whose regions were proposed by particular detectors favours those detectors. Information retrieval has known this about pooled judgements for decades [27, 28]. Real text found only by a new model counts as a false positive until someone adds it to the reference. Second, at the small data scale of a hobby project, training noise is large. Two runs of an identical recipe that differed only in seed disagreed as much as two different recipes did [29, 44].

1.3 Thesis and approach#

Our thesis is that for a comic-translation pipeline that edits what it detects, detectors should be compared at matched false-positive ceilings, against a reference corrected for pooling bias, with page-level and seed-level uncertainty reported. Judged this way, several conclusions that look settled at a fixed threshold turn out to be wrong, and the changes that actually help become visible.

We support the thesis with a case study that ran from August 26 to September 26, 2026. It produced the extension, a 57-page hand-reviewed detection reference, a training set of about 3,000 pages, and fifteen detector training runs, each scored the same way. The method is empirical throughout. Every design choice described below was fixed by a measurement on real pages, and several of those measurements overturned the hypothesis they were meant to confirm.

1.4 Contributions#

  1. A fully client-side comic translation system for Manifest V3 browsers (Section 3). It runs five stages in one shared offscreen document. The extension-origin document is exempt from CORS for image fetches, and credential-free fetching closes the confused-deputy risk that exemption opens. OpenCV.js runs in a sandboxed page because its runtime needs eval-class code generation.
  2. An evaluation protocol designed for destructive comic-translation pipelines (Section 4). It combines coverage-based recall, matched-spurious-ceiling sweeps, hand correction of a pooled reference, and page-bootstrap and multi-seed reporting.
  3. An empirical account of adapting a detector with noisy automatic labels (Section 5). Six YOLO runs isolate four candidate causes of a precision gap: the labelling rule, the starting checkpoint, hard negatives and training-set expansion. A fifth effect, seed variance, turned out to be as large as the differences between recipes.
  4. TextSeg (Section 6): a 12.8 MB text-probability-map detector trained with an ignore mask. It needs no Ultralytics code and is the first detector in this project to beat the baseline at the tight ceiling on every seed. We also describe how its map is reused to erase lettering without damaging balloon outlines.
  5. Measured decisions for the downstream stages (Section 7): script-specific OCR, column chunking for Japanese, balloon fill, a caption plate whose ink never falls below 4.5:1 contrast, and a single batched request per page for free-tier cloud translation.

1.5 Roadmap#

Section 2 places the work in the literature on comic understanding, scene-text detection and evaluation methodology. Section 3 describes the system. Section 4 defines the evaluation protocol, which the rest of the paper depends on. Section 5 follows the YOLO fine-tuning lineage and its diagnoses. Section 6 presents TextSeg, three training recipes and the analysis of what limits the tight ceiling. Section 7 covers OCR, erasing, rendering and translation. Section 8 discusses lessons, the development process, limitations and licensing. Section 9 concludes.


2.1 Comic and manga understanding#

Manga109 [1, 2] is the standard annotated manga corpus. Its bounding boxes for frames, faces, bodies and text supported early detection work [53]. A 2026 revision reports missing text regions, dialogue overlapping onomatopoeia, and under-segmented balloons in the original annotations [8]. These are the same defects we found in our own reference (Section 4.3). Magi [4] detects panels, text and characters jointly and orders dialogue for transcription. COO [5] targets comic onomatopoeia, which can be curved, truncated or split into parts, and treats their recognition and linking as open problems. Vivoli et al. [7] survey more than 300 comics papers and note that sound effects and text drawn on the art remain under-served. Del Gobbo and Herrera [6] argue that balloon detection alone cannot find manga text, since lettering often sits outside balloons, and they binarise text at the pixel level instead. Aramaki et al. [52] combined connected-component and region classifiers for manga text.

Hinami et al. [3] built the reference design for fully automatic manga translation. It has multimodal context-aware translation, a parallel corpus mined from published translations, and an end-to-end system with a server. Open-source tools [49, 50] use the same stage structure and depend on server-class models such as LaMa [24] for inpainting. Our work is complementary. We keep the stage structure but constrain the deployment to a browser, and we focus on how the detector should be evaluated under that pipeline’s error costs.

2.2 Text detection architectures#

Box regressors such as YOLO [16] and its NMS-free successors [17] suit browsers because their export is a fixed-shape tensor. The YOLO26 family used here is end-to-end, so no NMS has to be written in JavaScript. Transformer detectors such as RT-DETR [18, 19] are more accurate on comics. The model by ogkalu [47], trained on about 11,000 comic pages, reached sound-effect recall of 0.796 on our reference. At 168 MB and roughly 500 ms per page, though, we could use it only as a labelling teacher.

Segmentation-based scene-text detectors predict a per-pixel map and group it into regions afterwards. Examples include EAST [39], CRAFT [40] and DBNet [9]. DBNet shrinks each label polygon by an offset D=A(1−r2)/LD = A(1-r^2)/L and grows detected components back by the same offset, so touching instances stay apart. TextSeg adopts this label geometry and the combination of OHEM binary cross-entropy [13] and dice loss [14]. It sits on a MobileNetV3 backbone [11] with an FPN [12].

2.3 Recognition, inpainting and translation#

manga-ocr [46] is a ViT encoder [22] with a two-layer BERT decoder, in the TrOCR style [21], trained for Japanese manga. PP-OCR [20] combines a DB text-line detector with CTC recognisers [35] for Chinese and Korean. Telea’s fast-marching inpainting [23] is classical and cheap. LaMa [24] handles texture better but is too heavy for our target hardware. In-browser inference has moved from JavaScript tensor libraries such as TensorFlow.js [38] to WebAssembly and WebGPU runtimes. For on-device translation we use Gemma 2 [26] through WebLLM [25], which runs compiled WebGPU kernels at up to 80% of native speed.

2.4 Evaluation methodology#

We borrow three ideas from outside computer vision. First, pooling bias. Test collections judged only on items that participating systems retrieved favour those systems [27, 28], and a detection reference proposed by particular detectors has the same flaw. Second, variance. Bouthillier et al. [29] show that data sampling, initialisation and hyperparameters shift benchmark results enough to reorder methods, and Picard [44] finds that the seed alone moves results by amounts often reported as improvements. Third, operating points. Comparing classifiers at a fixed false-positive rate, rather than at a fixed threshold, is standard in ROC analysis [36]. We apply it with the false-positive rate set by the pipeline’s cost. For resampling we use the percentile bootstrap [30] over pages, the unit at which the reference was sampled.

Hard-negative mining has a long history in detection [13, 15]. Learning from partial or noisy labels is surveyed by Song et al. [34], and pseudo-labelling [33] is the usual way to scale up weak labels. Our automatic labeller is a pseudo-labeller over three independent detectors. Our ignore mask follows the partial-label idea: regions whose labels are too uncertain are left out of the loss rather than taught as background.

2.5 The gap#

Four things are missing from this literature. No study we know of evaluates comic text detectors at the operating point imposed by a pipeline that erases and overwrites detections. None measures how references pooled from detector proposals bias comparisons between comic detectors. None reports seed variance for detectors fine-tuned on a few thousand comic pages. None describes a complete translation pipeline running in the browser under Manifest V3. This paper addresses all four with one system and one dataset. The trade-off is breadth: our reference is small, and we report its uncertainty rather than hide it.


3. System: A Five-Stage Pipeline Inside a Manifest V3 Extension#

3.1 Where each stage can run#

A Manifest V3 extension [51] has four kinds of execution context, and each one rules out some stages. We checked each constraint before building on it.

  • Content scripts share the host page’s storage origin and are bound by its CORS rules. If the pipeline ran here, each site would get its own copy of the 1.4 GB translation model, and images from a CDN without CORS headers could not be read at all.
  • The service worker has no DOM and no WebGPU, which rules out on-device translation.
  • Extension pages run under script-src 'self' 'wasm-unsafe-eval'. This is enough for onnxruntime-web, but not for OpenCV.js. Its embind runtime builds every binding with new Function, which failed at initialisation under a simulated MV3 policy for both OpenCV 4.12 and 5.0. A control module loaded under the same guard succeeded, so the failure was OpenCV’s and not the test’s.
  • Sandboxed pages may use unsafe-eval but have no chrome.* APIs.

The design follows from these constraints (Figure 1). A content script finds the image and sends its URL. The whole pipeline runs in one offscreen document, created once per browser session, so every tab reuses the loaded models. Telea inpainting runs in a sandboxed child frame and exchanges pixel buffers with the offscreen document by transfer. An earlier version ran the pipeline in a hidden iframe per page. We verified that iframe design first: a Cache API marker written while it was framed by one origin could be read while it was framed by another, so the model cache was shared across sites. We later moved to the offscreen document so that models load once per browser session instead of once per tab.

flowchart TB
    IMG["img element on any site"] --> CS["Content script and floating control"]
    SW["Service worker: context menu, opens settings"] -.->|creates| OFF
    CS -->|image URL and settings| OFF["Offscreen document<br/>extension origin, cross-origin isolated<br/>fetches the image with credentials omitted"]
    OFF --> PIPE["Detect: TextSeg<br/>OCR: manga-ocr or PP-OCR<br/>Erase: balloon fill<br/>Translate: WebLLM or Gemini<br/>Render: Canvas 2D"]
    PIPE <-->|pixel buffers| SBX["Sandboxed page: OpenCV.js Telea"]
    PIPE -->|PNG data URL| CS

Figure 1. Runtime architecture. Only the content script touches the host page. The offscreen document owns every model.

3.2 Getting the pixels#

Comic images are almost always served from a CDN on a different origin. Our first design had the content script fetch the image and transfer an ImageBitmap. That failed on real sites in two ways. A bitmap from a non-CORS-clean image cannot be transferred at all (DataCloneError). Some image hosts send no Access-Control-Allow-Origin header at all; two we tested were nhentai.net and e-hentai.org. For those, no retry mode can help. Extension-origin documents with host permission are exempt from CORS, so the content script now sends only the URL, and the pipeline document fetches the bytes.

This required host_permissions: <all_urls>, which Chrome turns into its strongest install-time warning. It also created a security risk. The pipeline document is reachable from any page, so a credentialed fetch would let a hostile site use the extension to read a logged-in user’s private images on another origin. Every fetch therefore uses credentials: 'omit', which limits the extension to publicly readable images.

3.3 Triggering without the page’s cooperation#

On webtoons.com every image carries oncontextmenu="return false;". manhuagui.com cancels the context menu for the whole document. On either site, a right-click menu entry can never appear. We therefore added a floating control that relies on no event the page can cancel. It ranks on-screen images by how much of the viewport they cover and offers the top one. On the three sites we measured, that image covered 58%, 38% and 32% of the viewport, and the runner-up at most 6%. We rejected a cursor-following design because it depends on the pointer events those same sites suppress. It also cannot see images inside shadow roots: on reddit.com a document-level listener reports the shadow host, never the image.

Two features cover cases detection handles badly. The user can draw a box over missed text, and that region enters the pipeline where a detected box would, so no second code path exists. Tiled webtoon strips are detected and reported, not stitched. One measured strip was 291 tiles of 800×1000 pixels, 285,374 pixels tall, and a tile boundary ran through a speech balloon. Only 2 of the 291 tiles held pixels at any moment, because the viewer lazy-loads. We accepted per-tile translation as sufficient for this release.

3.4 Stage summary#

Table 1 lists the released configuration. All weights are pinned by revision and SHA-256 and downloaded at build time. The translation model is the exception: it downloads on first use and is cached in the Cache API.

Table 1. Released configuration (public release v0.0.1).

StageImplementationSizeNotes
DetectTextSeg, MobileNetV3-Large + FPN text map (Section 6)12.8 MBonnxruntime-web WASM, 4 threads
OCR, Japanesemanga-ocr [46], int8 encoder and decoder, column chunking117 MB + 9.4 MB line detector10/12 exact on the reference crops
OCR, ChinesePP-OCRv6 small det + small rec [20]31 MBCER 0.027 on horizontal text
OCR, Koreankorean PP-OCRv5 mobile rec13.4 MBshares the Chinese line detector
EraseBalloon fill, with Telea [23] in a sandbox as fallbackOpenCV.js 12.7 MBtext drawn on art is labelled instead (Section 7.3)
TranslateGemma 2 2B JPN, q4f16 [26] via WebLLM [25], or Gemini Flash-Lite1.4 GB (local)cloud model is opt-in, one request per page
RenderCanvas 2D typesetting, lobe-aware layout–caption plates with guaranteed contrast

Cross-origin isolation (COOP/COEP headers) gives the offscreen document SharedArrayBuffer and therefore four onnxruntime threads. With them, detection takes about 0.7 s per page and OCR about 0.4 s per region. Single-threaded, detection took 1.9 to 2.7 s, and Japanese OCR took about 1.15 s to encode each bubble plus 1.5 s for every 30 characters decoded. A twelve-bubble page needed roughly 30 s. Telea inpainting takes about 235 ms for a round trip to the sandbox, and layout and drawing take 83 ms together.


4. Evaluation Methodology#

4.1 The reference set#

We built detect_reference.json from 57 real pages. Thirty-seven came from pages where earlier sessions had logged failures, plus the OCR reference sets. Twenty were sampled from Danbooru by tag across the three source languages, with a minimum short side of 600 pixels and restricted to general and sensitive ratings. Danbooru shows a licence for each image, so it can be checked at fetch time. The file stores post IDs and boxes rather than pixels, so the set can be rebuilt without redistributing comic pages.

The ground truth was first proposed by two finders: the baseline YOLO26n at its 0.05 floor, and PP-OCR’s DB text detector. We inverted DB’s unclip in closed form so that its boxes bound the ink rather than the training margin. A reviewer then judged the clustered proposals on contact sheets and added missed regions by hand. Of the original 525 regions, 311 are bubbles, 72 captions and 142 sound effects, and 86 have dark backgrounds. After the corrections in Section 4.4 the final reference has 540 regions: 314 bubbles, 74 captions and 152 sound effects. The reviewer added 58 regions by hand that neither finder had found. One page was left out entirely: it carried dozens of repeated sound effects along its edges and could not be labelled exhaustively. A half-labelled page would have flattered every detector.

4.2 Metrics#

The pipeline crops each detected box for OCR. A box that is too loose costs a little inpainting, while a box that clips the lettering loses words. The primary measure is therefore coverage rather than IoU. For a ground-truth region gg and the set of detections D\mathcal{D}:

cov(g)=∣g∩⋃d∈Dd∣∣g∣,found(g)  ⟺  cov(g)≥0.5.\mathrm{cov}(g) = \frac{\left| g \cap \bigcup_{d \in \mathcal{D}} d \right|}{|g|}, \qquad \text{found}(g) \iff \mathrm{cov}(g) \ge 0.5 .

A threshold of 0.5 marks the point where OCR starts returning fragments instead of short readings. A detection is spurious when less than 0.25 of its area lies inside any ground-truth region. The cut is generous on purpose, because detector boxes are looser than the lettering inside them. The spurious rate is

Sm(t)=#{spurious detections of model m at threshold t}#{detections}.S_m(t) = \frac{\#\{\text{spurious detections of model } m \text{ at threshold } t\}}{\#\{\text{detections}\}} .

Recall is reported by slice, because the known gaps show up in slices rather than in the mean. The slices are kind (bubble, sound effect, caption), background, and language. A region counts as dark when at least 25% of its pixels fall below luma 80. We also report IoU>0.5 as the conventional localisation measure [41], together with merge and fragmentation counts.

4.3 Matched-ceiling comparison#

Detectors put their confidence on different scales, so a fixed threshold means something different for each model. We sweep thresholds from 0.05 to 0.95 on a 0.01 grid, reusing one inference pass per page. For each spurious ceiling cc we then take the threshold that maximises sound-effect recall without exceeding the ceiling:

tm∗(c)=arg⁡max⁡t : Sm(t)≤cRmsfx(t),t^{*}_m(c) = \arg\max_{t \,:\, S_m(t) \le c} R^{\text{sfx}}_m(t),

and we compare models on sound-effect, dark-background and overall recall at tm∗(c)t^{*}_m(c). The ceiling that decides shipping is c=0.9%c = 0.9\%, the baseline’s own rate at its default threshold. Looser ceilings (1.4% to 7%) show how the curves behave where the pipeline would tolerate more errors. The crossover is the lowest ceiling at which a candidate beats the baseline on both sound-effect and dark-background recall.

Two practical points matter here. First, spurious rate is not monotonic in threshold. Deduplication sees a different set of boxes at each threshold, so a lower threshold can have fewer spurious boxes. Second, the grid has to be fine where the operating points lie. An early grid had a 0.01 step only up to 0.65. Every TextSeg operating point sat above that, and one seed’s apparent failure at 0.9% was an artefact of the coarse grid (Section 6.3).

Important

At a fixed threshold, a detector that emits more boxes looks better on recall. At a matched spurious rate it has to earn its recall with the same false-positive budget. Every shipping decision in this paper uses the matched comparison. Fixed-threshold tables appear only for context.

4.4 Correcting the pooled reference#

The reference was proposed by the baseline detector and DB. Anything real that the baseline finds was therefore proposed, reviewed and labelled. Real text that only a new model finds is missing, so it counts against the new model. This is the pooling bias known from information retrieval [27, 28]. We corrected it the only safe way: by hand, one box at a time. The reviewer drew each new box around the ink, not from the candidate detector’s box, and never edited the JSON directly. The added regions were:

  • 7 regions from reviewing v4’s flagged boxes, 2 from v2’s and 1 from v3’s (525 to 535 regions);
  • 3 sound effects from TextSeg seed 2’s spurious boxes (538);
  • 1 sound effect from seed 0’s (539);
  • 1 sound effect from the recipe 2 gate review (540).

Seven of v4’s fifteen “spurious” boxes turned out to be unlabelled text. On the corrected reference, the genuine spurious rate at threshold 0.25 (535-region reference) is 0.6% for the baseline, 3.3% for v2 and 1.8% for v4. Once a comparison started, we never changed the reference in the middle of it. Every correction was followed by a re-sweep of all models.1

4.5 Uncertainty#

Two sources of noise dominate. Page sampling: the 0.9% gate rests on two or three boxes out of about 350. With about 150 sound-effect regions (147 at the 535-region stage, 152 in the final reference), the 95% interval on sound-effect recall is about ±0.075, which is as large as most differences between models. We therefore report a page bootstrap [30]: 1,000 resamples of the 57 pages, with the threshold re-chosen inside each resample, giving P(candidate beats the baseline on sfx and dark)P(\text{candidate beats the baseline on sfx and dark}), where beating means strictly higher recall on both. Training seeds: we report every seed separately and compare distributions, never single runs [29, 44].

4.6 Leakage guards#

The training, negative-mining and expansion sets are all sampled with the reference’s post IDs excluded (53 IDs were refused on the first training draw). The dataset builder and the training notebook both re-assert that no reference page is present. We also tested the guards by planting a reference page into a synthetic dataset directory and confirming that both build scripts refused it. When a review found informative false positives on reference pages, those boxes did not become training negatives. Only their categories informed the next mining round on fresh pages.


5. Adapting an Off-the-Shelf Detector: Six YOLO Runs#

5.1 Baseline and the two gaps#

At the 0.05 floor, the baseline found 0.974 of bubbles but only 0.415 of sound effects and 0.477 of dark-background regions. At its default threshold of 0.25 those two figures fell to 0.155 and 0.384. Both gaps are ones readers notice: sound effects are left untranslated, and dialogue on night scenes or dark panels is skipped. Both gaps are also missing from a mean score, because bubbles make up most of the reference.

5.2 A training set built without exhaustive labelling#

We sampled 2,925 Danbooru pages from outside the reference set, weighted toward the gaps. The weights were 1,170 sound-effect pages, 720 dark, 300 screentone, 540 Korean or Chinese, and 270 plain monochrome pages as a control, so that sound-effect recall could not be bought by forgetting bubbles. Hand-labelling at that scale was not possible, so labels came from three independent finders. The finders were proposers, and none of them was trusted as ground truth.

Two finders were not enough. On the original reference, the baseline and DB together proposed 90% of regions, but only 20% of the regions the reviewer had added by hand. Measured across all sound effects, 43 of 142 were found by neither. That mattered, because an unlabelled region is a negative in YOLO’s loss, so a set built from those two finders would have taught the model to suppress the very sound effects we wanted it to find. Adding ogkalu’s RT-DETR comic detector [47] as a third finder closed most of the hole. Its class for text outside balloons covers exactly the missing shape.

Clusters of proposals were then kept or dropped by a rule, autolabel.py, fitted against a hand-reviewed tier: 30 pages at first, 42 later, and 103 after the expansion described below. The first rule used finder confidence and box size. It labelled at precision 0.920 and recall 0.882 overall, and 0.766 / 0.735 on sound effects.

5.3 v1: recall solved, precision created#

Fine-tuning Kiuyha/Manga-Bubble-YOLO with Ultralytics [45] on this set took 20 epochs on a free Colab T4 (imgsz 1280, batch 8, AdamW, learning rate 0.002, 1.06 h). Two engineering fixes came first. Pages as large as 79 megapixels made data loading take 96 s per iteration until the export capped the long side at 1,600 pixels. Unthrottled fetching triggered HTTP 429 after about 400 pages and silently shrank the dataset until we added a delay and backoff. We also confirmed that the export path changed nothing: re-exporting the base checkpoint reproduced the baseline’s scores exactly, so any later change is due to training.

At threshold 0.25, v1 raised sound-effect recall from 0.155 to 0.465 and dark-background recall from 0.384 to 0.558. The spurious rate rose from 0.9% to 7.8%. The threshold sweep in Table 2 shows what that means.

Table 2. Fine-tune v1 against the baseline on the 525-region reference. Each row compares the two models at a similar spurious rate.

Spurious rateBaseline (threshold → sfx / dark / all)v1 (threshold → sfx / dark / all)
≈4.2%0.05 → 0.415 / 0.477 / 0.7920.40 → 0.366 / 0.465 / 0.701
≈3.4%0.10 → 0.296 / 0.453 / 0.7410.425 → 0.345 / 0.430 / 0.688
≈1.9%0.15 → 0.218 / 0.419 / 0.7050.475 → 0.275 / 0.349 / 0.640
≈0.8%0.25 → 0.155 / 0.384 / 0.6590.65 → 0.113 / 0.244 / 0.512

At equal false-positive cost, v1 loses on overall recall, bubble recall and IoU at every point. It wins on sound effects only in a narrow band between 2% and 3.4%, and by about 0.05 rather than the 0.31 the fixed-threshold view suggested. The largest improvement available without training was simply to lower the baseline’s own threshold. The baseline at 0.05, which users could already select, beat v1 at 0.40 on every slice at the same spurious rate. The model had learned the label noise as confident text. Because the noise did not sit at low confidence, no threshold could remove it.

5.4 v2: the label rule was reading the wrong variable#

We broke the rule’s false labels down by which finders had proposed each cluster, instead of tightening its thresholds. On the reviewed pages the base rates were far apart:

Finders that proposed the clusterReal text
Baseline YOLO, with or without others326 / 334 = 0.976
DB and RT-DETR only36 / 52 = 0.692
DB only10 / 28 = 0.357
RT-DETR only25 / 418 = 0.060

Fourteen of the old rule’s 22 false labels came from lone RT-DETR proposals. Inside that group RT-DETR’s confidence was inverted: the clusters the old rule kept scored 0.16 to 0.36 and were mostly art, while the real ones it dropped scored 0.05 to 0.14 and were large sound effects. DB proposals, for their part, arrive with no score at all, so the old score cut had silently thrown away every cluster DB found alone. The rebuilt rule looks at corroboration first and size second. Every YOLO-backed cluster is kept, DB-and-RT-DETR clusters need 2,500 px², DB-only clusters need 1,200 px², and anything else needs a lone finder at confidence 0.50 or higher. False labels fell from 22 to 6 while true positives rose from 254 to 256. A page-level bootstrap gave a precision gain of +0.057 (95% CI [+0.027, +0.088]). In split-half cross-validation, where the constants were fitted on 15 pages and scored on the other 15, the gain was +0.048 and positive in 99% of splits.

The reviewed set had included no screentone pages, and those were where the new rule changed the most. Twelve screentone pages were reviewed to cover them. The old rule’s precision on them was 0.849, against 0.920 elsewhere, and the new rule raised it to 0.925. Over all 42 reviewed pages, precision went from 0.899 to 0.961 and sound-effect precision from 0.695 to 0.878.

The relabelled set had 2,926 pages and 23,480 regions. We retrained v2 with the same budget as v1 (1.055 h against 1.056 h), so only the labels differed between the two runs. At threshold 0.25 the spurious rate fell from 7.8% to 4.6%, and every recall slice rose. That combination is what a label-quality fix looks like, as opposed to trading recall for precision. At the 0.9% ceiling, however, v2 still lost: it reached sound-effect 0.085 and dark 0.233 at threshold 0.65, against the baseline’s 0.155 and 0.384. Its crossover sat near 2.5%.

Fragmentation was a second symptom, not a second problem. The fine-tunes looked more fragmented: 13 fragmented regions against the baseline’s 4. The metric counted a region as fragmented whenever two detections each covered more than 5% of it, which also fires when a neighbouring box clips its edge. Only 5 of v2’s 13 were genuine splits. When the three models were matched on detection count rather than threshold (346, 346 and 344 boxes), the ordering reversed to 4, 3 and 2. The end-to-end head was intact, since duplicates before deduplication stayed at 0.047 to 0.058 per box, and removing the parent-drop rule changed nothing. One separable effect did appear: annotation granularity. The fine-tunes drew one box per vertical column where the baseline drew one per balloon, and each model reproduced the column-pair rate of its own training labels.

5.5 v3: the starting checkpoint is not the cause#

v1 and v2 both started from the manga-tuned checkpoint the extension already shipped. To test whether that starting point was holding the operating point back, and to move toward a model the project owned end to end, we retrained from Ultralytics’ generic COCO checkpoint on the same data. Two details show how easily a run’s meaning can shift. First, optimizer='auto' would silently have switched from AdamW to SGD at 40 epochs, so we pinned the optimizer. Second, Ultralytics 8.4 changed the export default, so the NMS-free head is exported only when nms=False is passed. Our shape assertion caught the resulting [1, 5, 33600] output before any scoring ran.

v3 needed about 27 epochs to reach the validation fitness v2 reached in 20. It ended almost on top of v2: 5.1% spurious at 0.25, with the same shape of curve. Two very different starting points gave one answer, so the confident false positives come from the data and labels, not from the initial weights.

5.6 v4: hard negatives are the first lever that moves the curve#

A later audit split the remaining label errors by page type. Screentone and hatching accounted for 7 of 15 false labels on the reviewed pages, marks backed by YOLO for 4, fragments for 2, and painterly colour pages for 2. We tested the hypothesis that dense, painterly art causes the false positives. Colourfulness [43] did not support it: precision was 0.940 on greyscale pages and 0.972 on colour pages.

So we mined the models’ own confusions. We ran the baseline, v2 and v3 over 700 fresh pages, weighted toward textures, effect marks and rendered art, and kept every box at threshold 0.25. That gave 7,795 candidates. The two text finders were used only to order the review. Tier A (no finder agrees) turned out to be 64% negatives, tier B (only RT-DETR agrees) 27%, and tier C (DB agrees) 1%. A and B were reviewed exhaustively and C was sampled. Every kept crop was then checked again at full size, because two crops that looked clean on the contact sheet turned out to hold clipped lettering.

The result was 158 confirmed negative crops on 123 pages. Of these, 94 were effect and costume marks, 63 rendered art, 1 a fragment, and 0 screentone. Two of the diagnosed categories could not be reached this way. A fragment always has lettering inside any honest window around it. Screentone never appeared as a confident false positive on fresh pages. The crops were packed onto grey canvases as background images, making up 5% of the training set, and v2’s recipe was otherwise left unchanged.

v4 was the first run to change the shape of the curve. At threshold 0.25 its spurious rate was 3.3% and every recall slice rose over v2. Genuine false positives at 0.25 fell from 15 to 8, and the ones that disappeared were exactly the marks and art the negatives targeted. Seven of v4’s remaining fifteen “spurious” boxes were unlabelled text (Section 4.4).

5.7 v5 and v4b: the seed moves results as much as the recipe#

A second mining round against v4 on 200 fresh pages found only 51 negatives. Most of what v4 now flagged was real text: in the DB-backed tier, 92 of 95 candidates were. The easy negatives were used up. We therefore expanded the true positives instead. Sixty-one fresh pages were reviewed with v4 as a fourth proposer: 587 regions were kept, 652 candidates rejected and 54 added by hand. The result was v5. v5 lost to v4 at every ceiling, and two cheap explanations failed. All 11 of v5’s spurious boxes were genuine errors. The expansion regions that only v4 had proposed improved the least, which is the opposite of what circularity would predict.

We then reran v5’s exact recipe with seed 1337, since v4 and v5 had both used seed 0 according to their published args.yaml. This replicate, v4b, settled the question (Table 3).

Table 3. Matched-ceiling recall (sfx / dark) on the 535-region reference. v5 and v4b are the same recipe with different seeds.

CeilingBaselinev2v3v4v5v4b
0.9%0.150 / 0.3840.136 / 0.2790.265 / 0.3840.265 / 0.3720.184 / 0.3260.327 / 0.442
1.9%0.259 / 0.4420.293 / 0.3950.333 / 0.4420.544 / 0.6630.463 / 0.5810.381 / 0.523
2.8%0.265 / 0.4420.449 / 0.5470.435 / 0.5470.612 / 0.7090.524 / 0.6740.503 / 0.628
Crossover–1.39%–0.96%1.32%0.81%

The two runs of the identical recipe disagree by as much as v4 and v5 do: the same-recipe spread in sound-effect recall is 0.143 at 0.9% and 0.082 at 1.9%. The order also changes with the ceiling. v4b is the best of the three at 0.9% and the worst at 1.9% and 2.8%. No single 20-epoch run is evidence about its recipe. We did not promote v4b, because doing so would repeat the mistake the replicate had just exposed. The baseline stayed in the extension, since the pipeline’s 0.9% tolerance sits inside this noise band.

gitGraph
    commit id: "baseline yolo26n"
    branch manga-yolo-lineage
    commit id: "v1 noisy labels"
    commit id: "v2 corroboration relabel"
    commit id: "v4 hard negatives"
    commit id: "v5 plus 61 pages"
    commit id: "v4b seed 1337"
    checkout main
    branch coco-start
    commit id: "v3 COCO checkpoint"
    checkout main
    branch textseg
    commit id: "R1 three seeds"
    commit id: "R2 fixed val, EMA"
    commit id: "R3 own negatives"
    checkout main
    merge textseg id: "ship TextSeg"

Figure 2. Detector lineage. Each commit is a trained and evaluated condition, and the three TextSeg commits are three seeds each.

Note

What the YOLO lineage established. The labels were a cause, and the corroboration rule fixed part of it. The starting checkpoint was not a cause. Hard negatives were the first lever to move the curve. At this data and compute scale, run-to-run variance was as large as the differences between recipes, so we compared distributions rather than single runs from then on.


6. TextSeg: A Text-Probability-Map Detector Trained With an Ignore Mask#

6.1 Why write a detector at all#

By this point the YOLO lineage had run into four limits. (i) The shipping test could not tell runs apart: at 0.9% on the 535-region reference, the baseline scored 2 false boxes of 347, v4 3 of 349, and v4b 3 of 372, so one box decided the outcome. (ii) 2,884 of the 2,987 training pages (96.6%) carried automatic labels, and the labeller finds only about 70% of sound effects. The rest were being taught as background, and Ultralytics offers no way to leave a region out of the loss. (iii) Every fine-tune scored about 0.61 on IoU>0.5 against 0.88 for the baseline, even as coverage rose, because the training boxes come from finders that draw a line, a text block or a whole balloon for the same lettering. (iv) Ultralytics states that its AGPL-3.0 licence covers models its training code produces, so every weight in the lineage inherited that provenance.

A per-pixel text map deals with (ii) to (iv) together. A column box and a block box paint mostly the same pixels. An ignore mask becomes one multiplication in a loss we write ourselves. The model, the loss and the training loop are this project’s own code, on an Apache-2.0 ImageNet backbone.

6.2 Model, labels and loss#

Model. A timm MobileNetV3-Large backbone [11] feeds features at strides 4, 8, 16 and 32 into an FPN [12] with 96 channels. Each level is smoothed to 24 channels, upsampled to stride 4, concatenated, and passed to a head that outputs one logit map. The head’s bias starts at −4-4 (σ(−4)≈0.02\sigma(-4) \approx 0.02), because nearly every pixel is background. Normalisation is built into the graph, so the ONNX file takes the same 0-to-1 input tensor, [1, 3, 1280, 1280], as the YOLO export did, and outputs [1, 1, 320, 320]. The file is 12.8 MB, and ONNX export drifts from PyTorch by less than 10−510^{-5}.

Labels. Each text box is painted shrunk in the DBNet manner [9], with every side moved in by

D=A (1−r2)L,r=0.4,D = \frac{A\,(1 - r^2)}{L}, \qquad r = 0.4,

where AA is the box area and LL its perimeter, so balloons that touch stay separate blobs. Every pixel of a training crop is one of four kinds. It is text if it lies inside a shrunk core. It is background if it lies elsewhere on the page. It is ignored if it lies inside a region the automatic labeller dropped even though DB had proposed it: such clusters are real text 36% to 69% of the time, so neither answer is safe to teach. Lone RT-DETR proposals and thin slivers are real only 6% and 1 in 29 of the time, so they stay background. The fourth kind is the ring between each core and its full box. At first we left the ring out of the loss. The smoke test showed that this let the model paint whole boxes, which the decoder then grew a second time, so the ring is now supervised as background, as DBNet does it. The full set holds 24,067 text regions and 1,848 ignore regions. Hard-negative pages enter whole, with everything except the confirmed windows ignored.

Loss. The loss is binary cross-entropy over masked pixels, keeping at most three hardest negatives per positive [13], plus a masked dice term [14]:

L=∑p∈Pℓp+∑n∈top3∣P∣(N)ℓn∣P∣+min⁡(∣N∣,3∣P∣)  +  1−2∑m y^ y+1∑m y^+∑m y+1,\mathcal{L} = \frac{\sum_{p \in P} \ell_p + \sum_{n \in \mathrm{top}_{3|P|}(N)} \ell_n}{|P| + \min(|N|, 3|P|)} \;+\; 1 - \frac{2\sum m\,\hat{y}\,y + 1}{\sum m\,\hat{y} + \sum m\,y + 1},

with mm the loss mask and ∣P∣|P| floored at 1, so that a crop with no text still pushes its background down. Decoding. The map is thresholded at 0.3, and each 4-connected component of at least 3 cells becomes a region, scored by its mean probability. Its box is grown back by the unshrink offset, which is found by solving the same fixed-point equation as training. A TypeScript port of the decoder in the extension gave identical boxes to the Python version on all 57 reference maps (1,136 boxes). The parity check fails if a single constant is changed.

Training. We trained on Kaggle [42] because a Kaggle run continues after the browser closes, so no machine has to stay on during training. The GPU was a T4 ×2, with AdamW at learning rate 10−310^{-3}, weight decay 10−410^{-4}, 768-pixel crops, batch 8, one warm-up epoch and cosine decay to 2%. Pages are cached once at a 1,536-pixel long side (2,987 pages plus 160 negative pages, 1.2 GB, no failed downloads), so each run reads from the cache instead of re-downloading. A CPU smoke test first checked that the model learns at all: after 90 s of training on 10 pages, pixel F1 on 2 unseen reviewed pages rose from 0.64 to 0.72.

6.3 Three recipes, nine runs#

Table 4. TextSeg matched-ceiling recall on the 540-region reference: range over seeds, with the mean. Every row, including the bootstrap row, is computed on the 540-region reference (1,000 page resamples; “beats” means strictly higher on both sound effects and dark backgrounds).

Ceiling, sliceBaselineR1: 80 ep, best-val epochR2: last-epoch EMAR2: val-peak EMAR3: R2 + own negatives
0.9% sfx0.1450.289–0.461 (0.390)0.289–0.414 (0.340)0.270–0.336 (0.307)0.375–0.441 (0.408)
0.9% dark0.3790.414–0.598 (0.525)0.425–0.517 (0.460)0.414–0.471 (0.441)0.506–0.563 (0.533)
1.9% sfx0.2500.487–0.553 (0.518)0.434–0.553 (0.480)0.572–0.592 (0.586)0.539–0.592 (0.568)
1.9% dark0.4370.563–0.678 (0.613)0.529–0.598 (0.563)0.655–0.678 (0.670)0.655–0.667 (0.663)
4.3% sfx0.3950.539–0.691 (0.629)0.605–0.658 (0.627)0.658–0.743 (0.700)0.717–0.743 (0.730)
4.3% dark0.4710.632–0.793 (0.686)0.655–0.724 (0.693)0.713–0.793 (0.751)0.782–0.805 (0.793)
PP(beats baseline at 0.9%), seeds 0/1/2–0.65 / 0.97 / 1.00–0.82 / 0.71 / 0.690.97 / 0.92 / 0.97

Recipe 1 used seeds 0, 1 and 2 for 80 epochs each, at about 82 s per epoch and 1 h 53 min to 2 h per run, and exported the epoch with the best validation F1. On the 535-region reference, as first reported, seeds 1 and 2 cleared the 0.9% gate and seed 0 did not. An independent audit of the raw logs overturned three readings of that result:

  1. The seeds were validated on different pages. The 25 validation pages were drawn with the training seed, so any two seeds shared only 4 to 8 of them. Validation F1 differed by 0.05 before training had even diverged. The apparent ordering in which “val F1 predicts the gate” (0.7085 < 0.7367 < 0.7544) could not be compared across seeds, and with three points a monotone order happens by chance one time in three.
  2. Seed 0’s failure was partly a grid artefact. Re-swept at 0.01 steps up to 0.95 on the 535-region reference, seed 0 does reach the 0.9% ceiling (threshold 0.81, 2 of 292 spurious). Its reported crossover of 1.14% could not be reproduced (0.97%).
  3. What set seed 0 apart was unsupervised padding. The loss masked out the grey letterbox, so the model’s output there was never trained. A component straddling the page edge survived decoding as a full-height sliver at x=0x = 0. These slivers were 46 of seed 0’s 83 spurious boxes at 0.25, against 15 for seed 1, 10 for seed 2 and none for the baseline. An unconstrained region is exactly where seeds should diverge.

The same audit tested a hypothesis that had been recorded as fact: that the training boxes were drawn larger than the reference boxes. The gap in size relative to the page is real. The median Japanese bubble box is 10.1 to 10.8 per mille of the page in training against 6.0 in the reference. But it reflects the pages, not the labelling. Boxes drawn by the same baseline detector at the same floor are 10.3 per mille on training pages and 5.9 on reference pages, because the training pages simply have larger balloons (median box side 194 px against 134 px). Most of TextSeg’s low IoU instead came from its own decoder (Section 6.5). The ignore mask, for its part, has no clean evidence either way. Recall above the labeller’s own sound-effect recall is not diagnostic, since v4 exceeds it too without any mask. An ablation would need many seeds per arm while the seed spread stays near 0.14.

Recipe 2 changed four things at once. It supervised the padding as background, fixed the validation pages across runs, cut training to 40 epochs (the best epochs had been 27 to 41), and kept an exponential moving average of the weights [31] with decay 0.999. All three seeds ran in a single 3 h 13 min Kaggle session. The page-edge slivers disappeared: zero at every threshold, against 46, 15 and 10 in recipe 1. The EMA’s validation F1 peaked at epoch 22 or 23 in every seed, and the last-epoch export had already lost recall. Exporting the EMA at its validation peak gave the tightest condition yet from 1.9% upward, with a seed spread of 0.020 in sound-effect recall against 0.066 for recipe 1. At 0.9%, however, nothing could be resolved: bootstrap probabilities were only 0.69 to 0.82.

The gate was the same few boxes in every seed. Across the three seeds, the spurious boxes at the 0.9% and 1.9% operating points came to 15 distinct boxes, and all were reviewed by hand. Thirteen were genuine errors: 8 were art, such as a “100” badge, a doodled face, a chibi and a shirt logo, and 5 were marks, such as scribbles, a ♪ and a running-man icon. One was text not meant for translation, and one was a real sound effect missing from the reference. All six boxes that set the 0.9% ceilings were genuine, and four of them fired in all three seeds. Three of the six sat on the reference set’s sketch pages, which carry only one or two regions each. The gate was therefore neither seed noise nor missing labels. It was a handful of glyph-like shapes that every run ranked above real text.

Recipe 3 targeted those shapes on fresh pages, never on the reference pages themselves. A Kaggle CPU notebook ran the recipe 2 models over 500 new pages and excluded 4,282 post IDs already used anywhere. The pages were weighted toward doodles, badges, scribbles, symbol marks and hatching. Punctuation-only bubbles were excluded from the weighting, because the reference counts ”?” and ”!” as text. Candidates scoring 0.5 or more went onto 17 review sheets (510 crops). The negatives were then checked again at full resolution, and this round kept 200 of them. Retraining with all three negative lists, from 3,263 training records including 301 negative pages, moved the 0.9% point on every seed. Each recipe 3 seed beats every recipe 2 seed on sound-effect and dark-background recall at the gate. The operating threshold fell from 0.74–0.76 to 0.70. The bootstrap probability of beating the baseline rose to 0.97, 0.92 and 0.97 for seeds 0, 1 and 2. The six gate boxes still fire, but at lower scores: the top one fell from 0.85 to 0.72.

Warning

What this does not show. For seeds 0 and 2 the 95% interval of the dark-background improvement at 0.9% touches zero, and for seed 1 it crosses zero. The gate still rests on two or three boxes out of about 390. We do not rank TextSeg against v4 or v4b: they differ in architecture, in labels and in seed variance, and no controlled comparison separates those factors.

6.4 Seed variance, measured directly#

Two recipes let us compare within-recipe spread to between-recipe differences. At 0.9%, recipe 1’s seeds spanned 0.29 to 0.46 in sound-effect recall. That range is wider than any difference between recipes measured in this project. Fixing the validation split and exporting a smooth EMA narrowed the spread above 1.9% threefold. Only the targeted negatives moved the 0.9% point on every seed. The practical rule we took from this: a recipe gets judged on three seeds and on page-bootstrap intervals, never on one run and never at a fixed threshold. Seed 2 of recipe 1 illustrates why. Its sound-effect recall at 0.25 was exactly 109/147, the same as seed 0’s, yet at matched ceilings they were the best and the worst of the three.

6.5 Box extent and the decoder#

The score threshold only filters whole components, so box extent is set entirely by the binarisation cut and the unshrink offset. That is why IoU stays flat as the threshold changes. We ran a decoder-only grid over BINARIZE ∈ {0.2, 0.3, 0.4, 0.5} and unshrink scales from 1.25 down to 0.25, on all three recipe 3 seeds, with no retraining. The unshrink scale traded coverage against IoU and nothing else. IoU>0.5 rose from 0.36 to 0.53 and then 0.67 as the scale fell from 1.0 to 0.75 and then 0.5, and every step cost coverage and matched-ceiling recall. The oversized boxes had been compensating for cores that sit poorly on the lettering. Since coverage is the primary metric, the decoder stayed at 0.3 with the full offset. Setting 0.2 with ×0.75 kept coverage (0.909 against 0.906) and added 0.13 IoU, so we noted it for erasing, where box extent matters.

6.6 Erasing with the map, not the box#

In the browser TextSeg found more text than YOLO, including ぬいっ, ガラ, スー, ビギッ and handwritten asides, but it erased worse. The erase stage relied on each box sitting inside its balloon so that Telea fills from clean interior, and TextSeg’s looser boxes crossed balloon outlines, which Telea smeared inward. The raw map could not replace the box, because it paints a shrunk core per text block rather than individual strokes. Erasing only the map’s cells left 13% to 36% of the lettering behind at every threshold and dilation we tried. Instead we grow each component’s core by a disk of 0.75 times its own unshrink offset. This rebuilds the box minus its corners, and the corners are where a rectangle crosses a round balloon. Section 7.3 gives the measurements.

TextSeg ships as the extension’s only detector at a default threshold of 0.60. This is looser than the 0.9% gate (threshold 0.70 for the deployed seed). We picked it by inspecting real pages in the development harness, where 0.5 to 0.6 added real lettering and almost no false boxes. It also sits near the model’s matched 1.9% point on the reference (0.61). The default was not derived from the gate, and calibrating it per language against the matched sweep is listed as open work (Section 8.3).


7. Downstream Stages: Decisions Fixed by Measurement#

The detector sets the operating point, and the downstream stages set what each region costs. This section reports the measurements behind each stage’s design. We give them briefly, because each one follows the same pattern: a plausible default, then a check on real data that confirmed it or overturned it.

7.1 Recognition: the script chooses the engine#

We sent Chinese and Korean text through manga-ocr to see how it would fail. It did not degrade gracefully: the output was unrelated to the input. Its vocabulary lacks 这, 种 and 门, and it contains only 191 of the 11,172 Hangul syllable blocks. Worse, it is confident when wrong: it assigned mean log-probabilities of −0.04 to −0.78 to garbled Chinese, mostly above the −0.5 cut our harness used to flag weak readings. That is why the source language is a user setting and not an automatic guess. A wrong guess produces fluent nonsense and gives no sign that it failed.

Table 5. OCR by script, measured on regions the detector found on real pages.

Script, setEngineExactCER
Japanese, 12 reference crops [46]manga-ocr, int8/int8 (111 MB)10/120.041
Japanese, same, without post-processingmanga-ocr6/12–
Japanese, 24 multi-column balloonswhole region16/240.047
chunks of ≤ 18 characters (released setting)21/240.004
Chinese, 16 horizontal regionsmanga-ocr1/160.420
PP-OCRv6 small det + small rec (31 MB)13/160.027
Chinese, 6 vertical columnsPP-OCR, fixed-pitch cells0/60.333
PP-OCR, ink-projection cells (Otsu [32])2/60.253
Korean, 67 regions on 11 pagesmanga-ocr0/671.080
PP-OCRv6 Chinese recogniser0/670.937
korean PP-OCRv5 mobile (13.4 MB)21/670.153

Several details in this table came from measurement rather than design. Porting manga-ocr’s post-processing, which folds half-width to full-width characters, doubled the number of exact readings. For Chinese, the choice of detector mattered more than the size of the recogniser: swapping in the older v5 line detector cost more than dropping from the medium recogniser to the small one. For Korean, two script-specific fixes mattered. Joining lines with a space lifted exact readings from 10 to 21. Line grouping had to compare box centres rather than overlap, because DB’s unclip grows every line box by about 22 px, so neighbouring lines always overlap. Vertical Korean never appeared: all 83 Korean regions on 11 pages were horizontal. We therefore report vertical Korean as untested rather than handled. For Japanese, manga-ocr squashes each crop to 224×224 pixels, so a four-column balloon leaves each character only a few pixels. Reading long regions in groups of at most 18 characters, using PP-OCR’s line detector to find the columns, fixed most multi-column misreads without changing the 12 standard crops.

7.2 Translation: the fallback that makes it reachable#

The on-device model was chosen from WebLLM’s 163 prebuilt models on the same 13 real OCR lines. On that set Qwen2.5-1.5B made a tenfold error in an amount of money and left a kanji untranslated in its English output. Gemma 2 2B JPN made neither error, at 1.49 GB against 0.88 GB. WebGPU is not universal, so we added a cloud fallback. We compared four providers on the constraints that decide the choice. CORS turned out not to matter, since the pipeline document is exempt anyway. The deciding factor was the free tier: only Gemini had one, and a fallback that needs a credit card would not help the users who lack WebGPU. DeepL also cannot be prompted, so it cannot follow the instruction to render sound effects as English sound effects.

Table 6. Meaning errors on 23 real OCR lines (12 Japanese, 5 Chinese, 6 Korean), same prompt and clean-up for both models.

JapaneseChineseKoreanof which SFXAllms/line
gemma-2-2b-jpn-it (local, WebGPU)6342 of 2122946*
gemini-3.5-flash-lite (cloud)0100 of 2184

*Measured on a laptop GPU that was held at 210 MHz of its 2,100 MHz by a power policy (0.6 TFLOPS against an expected 15 to 18). The latency figure describes that machine, not the model.

The local model changed a place name, dropped a chapter heading, turned a scream and a doorbell into unrelated exclamations, and reversed grammatical person in Korean. These are errors of content, not the tone-and-idiom limitations the project had expected when it started. Korean is the weakest language, as expected for a Japanese-tuned model. The cloud path required one more change to be usable at all. Sending one request per bubble got 7 of 23 lines rejected with HTTP 429 on the free tier. Sending one numbered request per page got none rejected and was 17 times faster per line. It was also slightly better, because the model saw the whole page’s dialogue. Three more practical facts about the cloud API: temperature: 0 did not make its output reproducible, a 96-token budget produced empty replies from models that spend tokens on reasoning, and the models-list endpoint was not a reliable sign that a model was available. When a selected target language is not supported by the local model, the extension stops and offers cloud translation. It never switches silently, because that would send the user’s text to a third party they had not chosen.

7.3 Erasing: follow the balloon#

We measured the erase stage on the reference pages with two quantities. Letters erased is the share of dark pixels (luma below 128) belonging to lettering that stays inside its reference box, and it should be high. Other ink is the number of dark pixels erased outside the lettering per page, which is damage to outlines and art, and it should be low.

Table 7. Erase masks on the 57 reference pages.

MaskLetters erasedOther ink erased (px/page)
YOLO box, padding 00.9852,656
TextSeg box0.96336,894
TextSeg shape, grow ×0.750.96712,403
YOLO box + balloon fill0.985892
TextSeg shape + balloon fill0.9849,038

Box padding was a case where the obvious choice was wrong. We started with 3 px of padding for anti-aliasing halos. Measured residual ink rose steadily with padding: 7.3% at 0 px, 19.3% at 3 px and 38.4% at 8 px. Padding pushes the ring Telea samples from off the clean balloon interior and onto the dark outline.

Balloon fill handles a case from a benchmark page used since the start of the project: a single line of speech running across two joined balloons. Both detectors return one region spanning the join, and Telea smears the outline between the balloons into an X. The fill samples the ground colour at the region’s text core and flood-fills pixels within 40 levels of it on every channel. The components that reach the core, with their holes filled, form the balloon. The letters are the holes, while the outline, including the one at the join, connects to the page outside and is left alone. The clipped region is then painted flat with the balloon’s colour. Handing a clipped mask back to Telea failed, because once the mask stops at the outline, the ring Telea fills from is the outline. A guard sends a region back to Telea when more than 5% of its ink is letter-like shapes the fill would leave behind. This guard came from a failure: the first version, which checked for ink lost from the middle of the region, also switched off the fill on the joined balloons it was meant for.

Text drawn on art is labelled, not erased. When a region gets no balloon fill and its surroundings fail the flat-field test below, its lettering is left in place. The translation goes in a small solid label beside it, at most 16 px and three lines. This removed the star-burst smears Telea produced on sound effects over artwork. One case remains open. A sound effect with a white outline over grey art looks like a balloon to the fill, so it gets painted as a white silhouette. We measured five hand-built separators, and none of them split outlined lettering from real balloons. The principled fix is a second output channel on TextSeg that predicts whether text sits inside a balloon, trained from the region kinds the labels already carry (14,528 bubble and 8,515 sound-effect regions in the automatic tier alone).

7.4 Rendering: contrast by construction#

Detector boxes around vertical Japanese are tall and narrow: 11 of 13 regions on a test page had aspect ratios between 0.15 and 0.86. In three cases the box was narrower than a single English word of its own translation at a 10 px font, so shrinking alone could never fit the text. The layout may therefore widen a line up to 2.5 times the box width, centre it, and flag the region as overflowed. It never truncates, because the translation is the only copy the reader gets. Growing a box into surrounding white space was rejected after a test in which the growth escaped a balloon into the page gutter, from 32 px to 203 px wide. A white balloon interior and a white page look the same to a brightness test. With balloon fill, text is now set inside the balloon’s own largest interior rectangles. A translation is split across the lobes of a joined balloon only when no part would come out smaller than the whole text would in the main lobe.

A caption plate goes behind text that has no balloon. The trigger is the share of the region, plus a 6 px margin, covered by its two dominant tones on the source page. Lettering on any flat field is two-tone by construction, and artwork is not. On 133 regions from 17 pages, a cut at 0.73 caught 5 of the 9 text-on-art regions with no false alarms among 124 bubbles. Looser cuts added false alarms faster than hits (0.89 gave 8 of 9 hits but 47 false alarms). Two other candidates failed. Ring uniformity separated perfectly on the first page and collapsed on held-out pages, flagging 24 of 57 regions with only about 5 correct, because a box that hugs its balloon puts the outline inside the ring. Measuring on the erased image instead of the source raised false alarms from 0 to 5, because Telea had already replaced the evidence with its own guess.

The plate takes the panel’s own tone, and the ink is black or white, whichever contrasts more. With relative luminance LL, the contrast against white ink is 1.05/(L+0.05)1.05/(L+0.05) and against black ink (L+0.05)/0.05(L+0.05)/0.05. The two are equal at

L∗=0.0525−0.05≈0.1791,contrast=0.05250.05≈4.58,L^{*} = \sqrt{0.0525} - 0.05 \approx 0.1791, \qquad \text{contrast} = \frac{\sqrt{0.0525}}{0.05} \approx 4.58 ,

so the better of the two choices never falls below the WCAG AA level of 4.5:1 [37] for any tone. Over all 136 regions the worst case is 8.57:1. Semi-transparency was the one gap in this guarantee. Blending with Telea’s output underneath dropped the worst case to 3.28:1 for grounds with luma between 89 and 141, so plates in that band are drawn opaque.


8. Discussion#

8.1 What generalises#

Judge detectors where the pipeline operates. In this study the fixed-threshold comparison was misleading every time it disagreed with the matched one. v1’s “+0.31 sound-effect recall” was mostly extra boxes. Seed 0 and seed 2 of recipe 1 looked identical on sound effects at 0.25 but were the worst and best at the gate. For a comic-translation pipeline that edits pixels, our experience suggests publishing the tolerated spurious rate and comparing models at that rate. We expect the same to hold for other pipelines that overwrite their detections, but this study does not test that.

A pooled reference is biased toward the detectors that proposed it. Half of v4’s apparent false positives were real text that the baseline had never proposed. The fix is hand review of every candidate’s flagged boxes, with the reference re-frozen before each comparison. As models improve, this matters more, because a better model finds more text that the reference does not yet contain. The 2026 revision of Manga109 [8] reports the same kinds of missing region.

Report variance before recipes. At our scale, the seed moved results as much as any recipe change did. Three seeds with a page bootstrap cost roughly three GPU-hours per recipe on a free T4, less than the Colab hours spent on single runs whose conclusions later reversed.

Diagnose labels by source, not by score. Which finders agreed on a region predicted real text far better than any finder’s confidence, and one finder’s confidence was inverted within its own group. Diagnostics grouped by source should come before any threshold tuning of an automatic labeller.

8.2 How this was built#

This was a side project, built in spare time over about a month, not a formal academic study. We still tried to document the measurements carefully enough that the numbers can be checked. LLM coding agents (Claude Code) did a large share of the work: they implemented the extension and the training code, planned and ran experiments, analysed and audited results, and drafted this write-up. The author reviewed the reported measurements, did the hand annotation and review, made the technical and product decisions, and edited the final text. That review was mostly reading and spot-checking, not an independent re-run of every number, so errors of the kind the audits below caught may remain.

Two habits mattered for the results. First, later sessions audited earlier ones against the raw training logs. That audit found the per-seed validation split, the coarse-grid artefact behind seed 0’s apparent gate failure, the unsupervised-padding slivers, and the evidence against the box-size hypothesis (Section 6.3). This was separate from an earlier tooling bug found during the v4b run itself: the sweep script’s --out option overwrote the saved rows of models not named on the command line, and those rows were restored from version control before any comparison was made. Second, agent-reported numbers were treated as claims until reproduced. Trained models were re-scored locally from the files downloaded back from Hugging Face or Kaggle. Wherever a run had logged its own scores, the local numbers matched them.

8.3 Limitations#

The reference has 57 pages, and the tight gate rests on two or three boxes, so we report bootstrap probabilities rather than claim significance. All pages come from one source, Danbooru, whose tag vocabulary shaped both sampling and slicing. The reviewing was done by one person without a second annotator, so labelling conventions (for example, excluding watermarks and printed cards from what needs translation) are one person’s. TextSeg’s ignore mask is untested by ablation, and its Japanese boxes remain about 1.5 times the reference area even with a tightened decoder. Browser detection differs slightly from the Python evaluation, because Chrome’s canvas resize is not PIL’s (per-region scores differ by up to 0.16). The translation comparison uses 23 lines and one annotator. On-device translation has not been re-verified since the pipeline moved to the offscreen document. Tiled webtoon strips are translated one tile at a time, not stitched.

8.4 Licensing and data ethics#

Weights trained with Ultralytics’ code are covered, in that company’s stated view, by AGPL-3.0, whatever licence tag a derivative carries. That was one reason for writing TextSeg. The released extension ships only TextSeg (Apache-2.0 backbone), manga-ocr and PP-OCR (Apache-2.0), OpenCV.js (Apache-2.0) and onnxruntime-web (MIT). Gemma is downloaded by the user’s browser under the Gemma Terms of Use. The repositories store post IDs and boxes, never pages. Every page is re-fetched from Danbooru, whose per-image licence can be checked at fetch time. We avoided sources such as scanned doujinshi, whose rights status is unclear. The public repository was assembled from scratch, with no shared git history, after a full-history secret scan of the private one.


9. Conclusion#

We set out to put a comic translator into the browser and found that its hardest problem was measurement. A pipeline that erases and rewrites what it detects has a false-positive budget. For our baseline detector that budget was 0.9%, and at that budget most apparent improvements disappeared. Matched-ceiling sweeps, a reference corrected for pooling bias, and seed-and-page uncertainty turned a sequence of confusing reversals into a consistent account. Labels mattered and the starting checkpoint did not. Hard negatives were the first lever that moved the curve. Single runs could not rank recipes. With these methods in place, TextSeg, a 12.8 MB detector we wrote ourselves with an ignore-aware loss, beats the baseline at the tight ceiling on every seed, on one 57-page reference. It roughly triples sound-effect recall (0.41 against 0.15) and raises dark-background recall from 0.38 to 0.53.

Three directions follow from what remains open. A balloon-versus-sound-effect channel in TextSeg would let the pipeline fill real balloons and label outlined sound effects, removing the last major erase failure. A larger, independently reviewed reference, extended by pooling every candidate’s proposals, would narrow the bootstrap intervals enough to test the ignore mask by ablation. Calibrating the default threshold per language, and moving detection and OCR to WebGPU, would make the matched-ceiling result the configuration users actually run. We think the protocol matters more than any single number in this paper. It was designed for, and tested on, one comic-translation pipeline; we expect it to be useful wherever a detector’s errors end up drawn on the page, but showing that would take other datasets and pipelines.


References#

[1] K. Aizawa, A. Fujimoto, A. Otsubo, T. Ogawa, Y. Matsui, K. Tsubota, H. Ikuta. “Building a Manga Dataset ‘Manga109’ with Annotations for Multimedia Applications.” IEEE MultiMedia, 27(2), 2020.

[2] Y. Matsui, K. Ito, Y. Aramaki, A. Fujimoto, T. Ogawa, T. Yamasaki, K. Aizawa. “Sketch-based Manga Retrieval using Manga109 Dataset.” Multimedia Tools and Applications, 76, 2017.

[3] R. Hinami, S. Ishiwatari, K. Yasuda, Y. Matsui. “Towards Fully Automated Manga Translation.” AAAI, 35(14):12998–13008, 2021. arXiv:2012.14271.

[4] R. Sachdeva, A. Zisserman. “The Manga Whisperer: Automatically Generating Transcriptions for Comics.” CVPR, 2024. arXiv:2401.10224.

[5] J. Baek, Y. Matsui, K. Aizawa. “COO: Comic Onomatopoeia Dataset for Recognizing Arbitrary or Truncated Texts.” ECCV, 2022. arXiv:2207.04675.

[6] J. Del Gobbo, R. Matuk Herrera. “Unconstrained Text Detection in Manga: A New Dataset and Baseline.” ECCV Workshops, 2020. arXiv:2009.04042.

[7] E. Vivoli et al. “One Missing Piece in Vision and Language: A Survey on Comics Understanding.” arXiv:2409.09502, 2024.

[8] “Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding.” ICML Culture × AI Workshop, 2026. arXiv:2605.21182.

[9] M. Liao, Z. Wan, C. Yao, K. Chen, X. Bai. “Real-time Scene Text Detection with Differentiable Binarization.” AAAI, 2020.

[10] mayocream. koharu (software). https://github.com/mayocream/koharu ↗

[11] A. Howard et al. “Searching for MobileNetV3.” ICCV, 2019.

[12] T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, S. Belongie. “Feature Pyramid Networks for Object Detection.” CVPR, 2017.

[13] A. Shrivastava, A. Gupta, R. Girshick. “Training Region-based Object Detectors with Online Hard Example Mining.” CVPR, 2016.

[14] F. Milletari, N. Navab, S.-A. Ahmadi. “V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation.” 3DV, 2016.

[15] P. Felzenszwalb, R. Girshick, D. McAllester, D. Ramanan. “Object Detection with Discriminatively Trained Part-Based Models.” IEEE TPAMI, 32(9), 2010.

[16] J. Redmon, S. Divvala, R. Girshick, A. Farhadi. “You Only Look Once: Unified, Real-Time Object Detection.” CVPR, 2016.

[17] A. Wang et al. “YOLOv10: Real-Time End-to-End Object Detection.” NeurIPS, 2024.

[18] Y. Zhao et al. “DETRs Beat YOLOs on Real-time Object Detection.” CVPR, 2024.

[19] W. Lv, Y. Zhao, Q. Chang, K. Huang, G. Wang, Y. Liu. “RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer.” arXiv:2407.17140, 2024.

[20] Y. Du et al. “PP-OCR: A Practical Ultra Lightweight OCR System.” arXiv:2009.09941, 2020.

[21] M. Li et al. “TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models.” AAAI, 2023.

[22] A. Dosovitskiy et al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.” ICLR, 2021.

[23] A. Telea. “An Image Inpainting Technique Based on the Fast Marching Method.” Journal of Graphics Tools, 9(1), 2004.

[24] R. Suvorov et al. “Resolution-robust Large Mask Inpainting with Fourier Convolutions.” WACV, 2022.

[25] C. F. Ruan et al. “WebLLM: A High-Performance In-Browser LLM Inference Engine.” arXiv:2412.15803, 2024.

[26] Gemma Team. “Gemma 2: Improving Open Language Models at a Practical Size.” arXiv:2408.00118, 2024.

[27] C. Buckley, D. Dimmick, I. Soboroff, E. Voorhees. “Bias and the Limits of Pooling for Large Collections.” Information Retrieval, 10:491–508, 2007.

[28] J. Zobel. “How Reliable Are the Results of Large-Scale Information Retrieval Experiments?” SIGIR, 1998.

[29] X. Bouthillier et al. “Accounting for Variance in Machine Learning Benchmarks.” MLSys, 2021.

[30] B. Efron. “Bootstrap Methods: Another Look at the Jackknife.” Annals of Statistics, 7(1), 1979.

[31] B. T. Polyak, A. B. Juditsky. “Acceleration of Stochastic Approximation by Averaging.” SIAM Journal on Control and Optimization, 30(4), 1992.

[32] N. Otsu. “A Threshold Selection Method from Gray-Level Histograms.” IEEE Transactions on Systems, Man, and Cybernetics, 9(1), 1979.

[33] D.-H. Lee. “Pseudo-Label: The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks.” ICML Workshop on Challenges in Representation Learning, 2013.

[34] H. Song, M. Kim, D. Park, Y. Shin, J.-G. Lee. “Learning from Noisy Labels with Deep Neural Networks: A Survey.” IEEE TNNLS, 2023.

[35] A. Graves, S. Fernández, F. Gomez, J. Schmidhuber. “Connectionist Temporal Classification.” ICML, 2006.

[36] T. Fawcett. “An Introduction to ROC Analysis.” Pattern Recognition Letters, 27(8), 2006.

[37] W3C. “Web Content Accessibility Guidelines (WCAG) 2.1.” W3C Recommendation, 2018.

[38] D. Smilkov et al. “TensorFlow.js: Machine Learning for the Web and Beyond.” SysML, 2019.

[39] X. Zhou et al. “EAST: An Efficient and Accurate Scene Text Detector.” CVPR, 2017.

[40] Y. Baek, B. Lee, D. Han, S. Yun, H. Lee. “Character Region Awareness for Text Detection.” CVPR, 2019.

[41] M. Everingham, L. Van Gool, C. Williams, J. Winn, A. Zisserman. “The PASCAL Visual Object Classes (VOC) Challenge.” IJCV, 88(2), 2010.

[42] Kaggle. “Notebooks” documentation (background “Save & Run All” execution, GPU quotas). https://www.kaggle.com/docs/notebooks ↗

[43] D. Hasler, S. Süsstrunk. “Measuring Colourfulness in Natural Images.” Proc. SPIE Human Vision and Electronic Imaging, 2003.

[44] D. Picard. “torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision.” arXiv:2109.08203, 2021.

[45] G. Jocher et al. Ultralytics YOLO (software, AGPL-3.0). https://github.com/ultralytics/ultralytics ↗

[46] kha-white. manga-ocr (software and model, Apache-2.0). https://github.com/kha-white/manga-ocr ↗

[47] ogkalu. comic-text-and-bubble-detector (RT-DETR-v2 model, Apache-2.0). Hugging Face.

[48] Kiuyha. Manga-Bubble-YOLO (YOLO26n model). Hugging Face.

[49] ogkalu. comic-translate (software). https://github.com/ogkalu2/comic-translate ↗

[50] zyddnys. manga-image-translator (software). https://github.com/zyddnys/manga-image-translator ↗

[51] Google. “Manifest V3” and “chrome.offscreen” developer documentation. https://developer.chrome.com/docs/extensions ↗

[52] Y. Aramaki, Y. Matsui, T. Yamasaki, K. Aizawa. “Text Detection in Manga by Combining Connected-Component-Based and Region-Based Classifications.” ICIP, 2016.

[53] T. Ogawa, A. Otsubo, R. Narita, Y. Matsui, T. Yamasaki, K. Aizawa. “Object Detection for Comics using Manga109 Annotations.” arXiv:1803.08670, 2018.


Appendix A. Artefacts and Reproducibility#

Released artefacts
ArtefactLocation
Extension source, MIT, TextSeg versionvermilion10/comic-translator (GitHub), release v0.0.1
TextSeg weightsHugging Face vermilion10/manga-textseg
YOLO fine-tunes v1, v2, v4, v5, v4bHugging Face vermilion10/manga-bubble-yolo-finetune (tags)
YOLO v3 (COCO start)Hugging Face vermilion10/manga-text-yolo26, tag v3
Detection reference (57 pages, 540 regions)ml/eval/detect_reference.json (post IDs and boxes)
OCR referencesml/eval/{zh,ko,ja_page}_reference.json
Sweep and scoringml/scripts/sweep_detect.py, ml/eval/score_detect.py
Parity checks (Python ↔ TypeScript)extension/scripts/check-text-map.mjs, check-balloon.mjs
Training budgets
RunHardwareEpochsWall time
YOLO v1, v2, v4, v5, v4bColab T4201.06–1.12 h each
YOLO v3Colab T4402.19 h
TextSeg recipe 1, per seedKaggle T4 ×2801.88–2.00 h
TextSeg recipe 2, three seedsKaggle T4 ×2403.22 h total
TextSeg recipe 3, three seedsKaggle T4 ×2403.04 h total
Timeline (2026)
  • Aug 26–28. Proposal, extension scaffold, then detect, OCR, erase, translate and render wired stage by stage.
  • Aug 28–31. Page integration, host permissions, options, Chinese and Korean OCR, deduplication, floating control, manual boxes, caption plates, and the 57-page detection reference.
  • Sep 2–3. Training set, v1, the threshold sweep, the corroboration relabel, v2, the fragmentation diagnosis.
  • Sep 10–12. Cloud fallback, vertical Chinese OCR, release audit, v3, false-positive taxonomy, hard-negative mining.
  • Sep 13–22. v4, reference correction, second mining round, expansion (v5), seed replicate (v4b).
  • Sep 23–26. TextSeg design and three recipes on Kaggle, map-shaped erasing, balloon fill, Japanese chunking, offscreen pipeline, public release.

Footnotes#

  1. The reference grew from 525 to 540 regions as corrections were added, and every comparison was made on the reference as it stood at the time. Each detection table names its reference (Table 2: 525, Table 3: 535, Table 4: 540), and all rows in a table share it. Numbers quoted in the text say which reference they come from when it is not the final 540-region one. ↩

Detecting Comic Lettering Under a False-Positive Budget: A Measurement-Driven Study of an In-Browser Comic Translator
https://pure.vermilion10.dev/blog/comic-translator-paper
Author vermilion10
Published at September 26, 2026