<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet href="/scripts/pretty-feed-v3.xsl" type="text/xsl"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:h="http://www.w3.org/TR/html4/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>vermilion10</title><description>Web developer and reverse engineer based in Bandung, Indonesia.</description><link>https://pure.vermilion10.dev</link><item><title>Detecting Comic Lettering Under a False-Positive Budget: A Measurement-Driven Study of an In-Browser Comic Translator</title><link>https://pure.vermilion10.dev/blog/comic-translator-paper</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/comic-translator-paper</guid><description>A fully client-side manga, manhua and manhwa translator for Chrome, and what 15 detector training runs taught us about evaluating text detection at the operating point a destructive pipeline can actually tolerate.</description><pubDate>Sat, 26 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;vermilion10&lt;/strong&gt; · Personal project write-up · September 2026&lt;/p&gt;
&lt;p&gt;::github{repo=&quot;vermilion10/comic-translator&quot;}&lt;/p&gt;
&lt;h2&gt;Abstract&lt;/h2&gt;
&lt;p&gt;Automatic comic translation chains five stages: text detection, recognition, erasure, translation and typesetting. Each stage acts on every region the detector returns, so a false detection is not a harmless extra box. It erases part of the artwork and prints an invented translation on it. We built a comic translator that runs entirely inside a Chrome extension. It covers Japanese, Chinese and Korean sources, and every model runs client-side. We use it to study how text detectors should be evaluated when false positives carry this cost. Seen in that light, the usual evaluation is misleading. Our first fine-tuned detector tripled sound-effect recall at the default threshold, yet at the pipeline&apos;s tolerated false-positive rate of 0.9% it was worse than the model it was meant to replace. We therefore compare detectors at matched spurious-detection ceilings. We also correct a reference set whose ground truth had been pooled from the proposals of the baseline detector and one text finder, and we resample pages and training seeds before believing a ranking. Across six YOLO fine-tunes we found that fixing the labels and mining hard negatives moved the operating curve, while changing the starting checkpoint did not. We also found that seed-to-seed variance was as large as the difference between recipes. We then trained our own detector, TextSeg: a MobileNetV3 and FPN text-probability map with an ignore-aware DBNet-style loss. Combined with negatives mined against TextSeg itself, it beats the baseline detector at the 0.9% ceiling on all three seeds, with page-bootstrap probabilities of 0.92 to 0.97. Its sound-effect recall roughly triples (0.41 against 0.15). Around the detector we report measured choices for per-script OCR, balloon-aware erasing, contrast-guaranteed caption plates and free-tier cloud translation. We designed the evaluation protocol for comic-translation pipelines that erase what they detect, and we expect it to be useful for similar pipelines, though we have tested it only on this one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Keywords:&lt;/strong&gt; comic text detection, manga translation, operating-point evaluation, pooling bias, hard-negative mining, in-browser inference&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;1. Introduction&lt;/h2&gt;
&lt;h3&gt;1.1 The problem&lt;/h3&gt;
&lt;p&gt;A reader who opens a Japanese manga, a Chinese manhua or a Korean webtoon in a browser sees text that most of them cannot read. Tools that translate these images already exist. zyddnys/manga-image-translator [50], comic-translate [49] and koharu [10] follow the same five-stage pattern: detect the lettering, recognise it, erase it, translate it and typeset the translation back into the art. Nearly all of them run as desktop applications or behind a server, so the reader has to save a page, open another program and come back. Hinami et al. [3] built a complete research system around the same pipeline and showed that translation quality depends on the image context as well as the text.&lt;/p&gt;
&lt;p&gt;We wanted the same pipeline with no step between reading and translating. That meant a browser extension that works on the page the reader is already viewing. It had to keep the text on the device by default and send nothing to a server unless the reader chose a cloud translator. Every stage therefore runs client-side: ONNX models on onnxruntime-web, OpenCV.js in a sandboxed page, and a quantised language model on WebGPU through WebLLM [25].&lt;/p&gt;
&lt;h3&gt;1.2 Why detection has to be judged at an operating point&lt;/h3&gt;
&lt;p&gt;The detector decides what the rest of the pipeline touches. Each box it emits is cropped for OCR, erased by inpainting, translated and overwritten with English. A missed sound effect stays untranslated, which the reader barely notices. A false box on a face, a heart or a speed line erases part of the drawing and prints an invented translation on it, and the reader notices that at once. So the two error types cost very different amounts. The extension first shipped with a YOLO26n model fine-tuned for manga bubbles [48]. We call it &lt;strong&gt;the baseline&lt;/strong&gt; throughout, because every candidate in this paper is compared against it; TextSeg (Section 6) later replaced it in the extension. On the original 525-region reference, the baseline produced 3 spurious boxes out of 346 detections (0.9%) at its default threshold. We take that rate as the pipeline&apos;s tolerance.&lt;/p&gt;
&lt;p&gt;Detection papers for comics usually report average precision, F-measure or pixel-level scores over a fixed test set [4, 6, 53], and model comparisons are usually made at one threshold. That is the right summary when every error costs the same. It is the wrong one here, and our own work shows why. Our first fine-tune raised sound-effect recall at threshold 0.25 from 0.155 to 0.465. It also raised the spurious rate from 0.9% to 7.8%. At a matched spurious rate of 0.9% the fine-tune lost to the baseline on every slice. Most of the gain had come from emitting 462 boxes instead of 346.&lt;/p&gt;
&lt;p&gt;Two more problems sit under any such comparison. First, a reference set whose regions were proposed by particular detectors favours those detectors. Information retrieval has known this about pooled judgements for decades [27, 28]. Real text found only by a new model counts as a false positive until someone adds it to the reference. Second, at the small data scale of a hobby project, training noise is large. Two runs of an identical recipe that differed only in seed disagreed as much as two different recipes did [29, 44].&lt;/p&gt;
&lt;h3&gt;1.3 Thesis and approach&lt;/h3&gt;
&lt;p&gt;Our thesis is that &lt;strong&gt;for a comic-translation pipeline that edits what it detects, detectors should be compared at matched false-positive ceilings, against a reference corrected for pooling bias, with page-level and seed-level uncertainty reported.&lt;/strong&gt; Judged this way, several conclusions that look settled at a fixed threshold turn out to be wrong, and the changes that actually help become visible.&lt;/p&gt;
&lt;p&gt;We support the thesis with a case study that ran from August 26 to September 26, 2026. It produced the extension, a 57-page hand-reviewed detection reference, a training set of about 3,000 pages, and fifteen detector training runs, each scored the same way. The method is empirical throughout. Every design choice described below was fixed by a measurement on real pages, and several of those measurements overturned the hypothesis they were meant to confirm.&lt;/p&gt;
&lt;h3&gt;1.4 Contributions&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;A fully client-side comic translation system&lt;/strong&gt; for Manifest V3 browsers (Section 3). It runs five stages in one shared offscreen document. The extension-origin document is exempt from CORS for image fetches, and credential-free fetching closes the confused-deputy risk that exemption opens. OpenCV.js runs in a sandboxed page because its runtime needs &lt;code&gt;eval&lt;/code&gt;-class code generation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;An evaluation protocol designed for destructive comic-translation pipelines&lt;/strong&gt; (Section 4). It combines coverage-based recall, matched-spurious-ceiling sweeps, hand correction of a pooled reference, and page-bootstrap and multi-seed reporting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;An empirical account of adapting a detector with noisy automatic labels&lt;/strong&gt; (Section 5). Six YOLO runs isolate four candidate causes of a precision gap: the labelling rule, the starting checkpoint, hard negatives and training-set expansion. A fifth effect, seed variance, turned out to be as large as the differences between recipes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TextSeg&lt;/strong&gt; (Section 6): a 12.8 MB text-probability-map detector trained with an ignore mask. It needs no Ultralytics code and is the first detector in this project to beat the baseline at the tight ceiling on every seed. We also describe how its map is reused to erase lettering without damaging balloon outlines.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Measured decisions for the downstream stages&lt;/strong&gt; (Section 7): script-specific OCR, column chunking for Japanese, balloon fill, a caption plate whose ink never falls below 4.5:1 contrast, and a single batched request per page for free-tier cloud translation.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;1.5 Roadmap&lt;/h3&gt;
&lt;p&gt;Section 2 places the work in the literature on comic understanding, scene-text detection and evaluation methodology. Section 3 describes the system. Section 4 defines the evaluation protocol, which the rest of the paper depends on. Section 5 follows the YOLO fine-tuning lineage and its diagnoses. Section 6 presents TextSeg, three training recipes and the analysis of what limits the tight ceiling. Section 7 covers OCR, erasing, rendering and translation. Section 8 discusses lessons, the development process, limitations and licensing. Section 9 concludes.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;2. Related Work&lt;/h2&gt;
&lt;h3&gt;2.1 Comic and manga understanding&lt;/h3&gt;
&lt;p&gt;Manga109 [1, 2] is the standard annotated manga corpus. Its bounding boxes for frames, faces, bodies and text supported early detection work [53]. A 2026 revision reports missing text regions, dialogue overlapping onomatopoeia, and under-segmented balloons in the original annotations [8]. These are the same defects we found in our own reference (Section 4.3). Magi [4] detects panels, text and characters jointly and orders dialogue for transcription. COO [5] targets comic onomatopoeia, which can be curved, truncated or split into parts, and treats their recognition and linking as open problems. Vivoli et al. [7] survey more than 300 comics papers and note that sound effects and text drawn on the art remain under-served. Del Gobbo and Herrera [6] argue that balloon detection alone cannot find manga text, since lettering often sits outside balloons, and they binarise text at the pixel level instead. Aramaki et al. [52] combined connected-component and region classifiers for manga text.&lt;/p&gt;
&lt;p&gt;Hinami et al. [3] built the reference design for fully automatic manga translation. It has multimodal context-aware translation, a parallel corpus mined from published translations, and an end-to-end system with a server. Open-source tools [49, 50] use the same stage structure and depend on server-class models such as LaMa [24] for inpainting. Our work is complementary. We keep the stage structure but constrain the deployment to a browser, and we focus on how the detector should be evaluated under that pipeline&apos;s error costs.&lt;/p&gt;
&lt;h3&gt;2.2 Text detection architectures&lt;/h3&gt;
&lt;p&gt;Box regressors such as YOLO [16] and its NMS-free successors [17] suit browsers because their export is a fixed-shape tensor. The YOLO26 family used here is end-to-end, so no NMS has to be written in JavaScript. Transformer detectors such as RT-DETR [18, 19] are more accurate on comics. The model by ogkalu [47], trained on about 11,000 comic pages, reached sound-effect recall of 0.796 on our reference. At 168 MB and roughly 500 ms per page, though, we could use it only as a labelling teacher.&lt;/p&gt;
&lt;p&gt;Segmentation-based scene-text detectors predict a per-pixel map and group it into regions afterwards. Examples include EAST [39], CRAFT [40] and DBNet [9]. DBNet shrinks each label polygon by an offset $D = A(1-r^2)/L$ and grows detected components back by the same offset, so touching instances stay apart. TextSeg adopts this label geometry and the combination of OHEM binary cross-entropy [13] and dice loss [14]. It sits on a MobileNetV3 backbone [11] with an FPN [12].&lt;/p&gt;
&lt;h3&gt;2.3 Recognition, inpainting and translation&lt;/h3&gt;
&lt;p&gt;manga-ocr [46] is a ViT encoder [22] with a two-layer BERT decoder, in the TrOCR style [21], trained for Japanese manga. PP-OCR [20] combines a DB text-line detector with CTC recognisers [35] for Chinese and Korean. Telea&apos;s fast-marching inpainting [23] is classical and cheap. LaMa [24] handles texture better but is too heavy for our target hardware. In-browser inference has moved from JavaScript tensor libraries such as TensorFlow.js [38] to WebAssembly and WebGPU runtimes. For on-device translation we use Gemma 2 [26] through WebLLM [25], which runs compiled WebGPU kernels at up to 80% of native speed.&lt;/p&gt;
&lt;h3&gt;2.4 Evaluation methodology&lt;/h3&gt;
&lt;p&gt;We borrow three ideas from outside computer vision. First, pooling bias. Test collections judged only on items that participating systems retrieved favour those systems [27, 28], and a detection reference proposed by particular detectors has the same flaw. Second, variance. Bouthillier et al. [29] show that data sampling, initialisation and hyperparameters shift benchmark results enough to reorder methods, and Picard [44] finds that the seed alone moves results by amounts often reported as improvements. Third, operating points. Comparing classifiers at a fixed false-positive rate, rather than at a fixed threshold, is standard in ROC analysis [36]. We apply it with the false-positive rate set by the pipeline&apos;s cost. For resampling we use the percentile bootstrap [30] over pages, the unit at which the reference was sampled.&lt;/p&gt;
&lt;p&gt;Hard-negative mining has a long history in detection [13, 15]. Learning from partial or noisy labels is surveyed by Song et al. [34], and pseudo-labelling [33] is the usual way to scale up weak labels. Our automatic labeller is a pseudo-labeller over three independent detectors. Our ignore mask follows the partial-label idea: regions whose labels are too uncertain are left out of the loss rather than taught as background.&lt;/p&gt;
&lt;h3&gt;2.5 The gap&lt;/h3&gt;
&lt;p&gt;Four things are missing from this literature. No study we know of evaluates comic text detectors at the operating point imposed by a pipeline that erases and overwrites detections. None measures how references pooled from detector proposals bias comparisons between comic detectors. None reports seed variance for detectors fine-tuned on a few thousand comic pages. None describes a complete translation pipeline running in the browser under Manifest V3. This paper addresses all four with one system and one dataset. The trade-off is breadth: our reference is small, and we report its uncertainty rather than hide it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;3. System: A Five-Stage Pipeline Inside a Manifest V3 Extension&lt;/h2&gt;
&lt;h3&gt;3.1 Where each stage can run&lt;/h3&gt;
&lt;p&gt;A Manifest V3 extension [51] has four kinds of execution context, and each one rules out some stages. We checked each constraint before building on it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Content scripts&lt;/strong&gt; share the host page&apos;s storage origin and are bound by its CORS rules. If the pipeline ran here, each site would get its own copy of the 1.4 GB translation model, and images from a CDN without CORS headers could not be read at all.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The service worker&lt;/strong&gt; has no DOM and no WebGPU, which rules out on-device translation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Extension pages&lt;/strong&gt; run under &lt;code&gt;script-src &apos;self&apos; &apos;wasm-unsafe-eval&apos;&lt;/code&gt;. This is enough for onnxruntime-web, but not for OpenCV.js. Its embind runtime builds every binding with &lt;code&gt;new Function&lt;/code&gt;, which failed at initialisation under a simulated MV3 policy for both OpenCV 4.12 and 5.0. A control module loaded under the same guard succeeded, so the failure was OpenCV&apos;s and not the test&apos;s.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sandboxed pages&lt;/strong&gt; may use &lt;code&gt;unsafe-eval&lt;/code&gt; but have no &lt;code&gt;chrome.*&lt;/code&gt; APIs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The design follows from these constraints (Figure 1). A content script finds the image and sends its URL. The whole pipeline runs in one offscreen document, created once per browser session, so every tab reuses the loaded models. Telea inpainting runs in a sandboxed child frame and exchanges pixel buffers with the offscreen document by transfer. An earlier version ran the pipeline in a hidden iframe per page. We verified that iframe design first: a Cache API marker written while it was framed by one origin could be read while it was framed by another, so the model cache was shared across sites. We later moved to the offscreen document so that models load once per browser session instead of once per tab.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;flowchart TB
    IMG[&quot;img element on any site&quot;] --&gt; CS[&quot;Content script and floating control&quot;]
    SW[&quot;Service worker: context menu, opens settings&quot;] -.-&gt;|creates| OFF
    CS --&gt;|image URL and settings| OFF[&quot;Offscreen document&amp;#x3C;br/&gt;extension origin, cross-origin isolated&amp;#x3C;br/&gt;fetches the image with credentials omitted&quot;]
    OFF --&gt; PIPE[&quot;Detect: TextSeg&amp;#x3C;br/&gt;OCR: manga-ocr or PP-OCR&amp;#x3C;br/&gt;Erase: balloon fill&amp;#x3C;br/&gt;Translate: WebLLM or Gemini&amp;#x3C;br/&gt;Render: Canvas 2D&quot;]
    PIPE &amp;#x3C;--&gt;|pixel buffers| SBX[&quot;Sandboxed page: OpenCV.js Telea&quot;]
    PIPE --&gt;|PNG data URL| CS
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Figure 1. Runtime architecture. Only the content script touches the host page. The offscreen document owns every model.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;3.2 Getting the pixels&lt;/h3&gt;
&lt;p&gt;Comic images are almost always served from a CDN on a different origin. Our first design had the content script fetch the image and transfer an &lt;code&gt;ImageBitmap&lt;/code&gt;. That failed on real sites in two ways. A bitmap from a non-CORS-clean image cannot be transferred at all (&lt;code&gt;DataCloneError&lt;/code&gt;). Some image hosts send no &lt;code&gt;Access-Control-Allow-Origin&lt;/code&gt; header at all; two we tested were ||nhentai.net|| and ||e-hentai.org||. For those, no retry mode can help. Extension-origin documents with host permission are exempt from CORS, so the content script now sends only the URL, and the pipeline document fetches the bytes.&lt;/p&gt;
&lt;p&gt;This required &lt;code&gt;host_permissions: &amp;#x3C;all_urls&gt;&lt;/code&gt;, which Chrome turns into its strongest install-time warning. It also created a security risk. The pipeline document is reachable from any page, so a credentialed fetch would let a hostile site use the extension to read a logged-in user&apos;s private images on another origin. Every fetch therefore uses &lt;code&gt;credentials: &apos;omit&apos;&lt;/code&gt;, which limits the extension to publicly readable images.&lt;/p&gt;
&lt;h3&gt;3.3 Triggering without the page&apos;s cooperation&lt;/h3&gt;
&lt;p&gt;On webtoons.com every image carries &lt;code&gt;oncontextmenu=&quot;return false;&quot;&lt;/code&gt;. manhuagui.com cancels the context menu for the whole document. On either site, a right-click menu entry can never appear. We therefore added a floating control that relies on no event the page can cancel. It ranks on-screen images by how much of the viewport they cover and offers the top one. On the three sites we measured, that image covered 58%, 38% and 32% of the viewport, and the runner-up at most 6%. We rejected a cursor-following design because it depends on the pointer events those same sites suppress. It also cannot see images inside shadow roots: on reddit.com a document-level listener reports the shadow host, never the image.&lt;/p&gt;
&lt;p&gt;Two features cover cases detection handles badly. The user can draw a box over missed text, and that region enters the pipeline where a detected box would, so no second code path exists. Tiled webtoon strips are detected and reported, not stitched. One measured strip was 291 tiles of 800×1000 pixels, 285,374 pixels tall, and a tile boundary ran through a speech balloon. Only 2 of the 291 tiles held pixels at any moment, because the viewer lazy-loads. We accepted per-tile translation as sufficient for this release.&lt;/p&gt;
&lt;h3&gt;3.4 Stage summary&lt;/h3&gt;
&lt;p&gt;Table 1 lists the released configuration. All weights are pinned by revision and SHA-256 and downloaded at build time. The translation model is the exception: it downloads on first use and is cached in the Cache API.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Table 1. Released configuration (public release v0.0.1).&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| Stage | Implementation | Size | Notes |
|---|---|---|---|
| Detect | TextSeg, MobileNetV3-Large + FPN text map (Section 6) | 12.8 MB | onnxruntime-web WASM, 4 threads |
| OCR, Japanese | manga-ocr [46], int8 encoder and decoder, column chunking | 117 MB + 9.4 MB line detector | 10/12 exact on the reference crops |
| OCR, Chinese | PP-OCRv6 small det + small rec [20] | 31 MB | CER 0.027 on horizontal text |
| OCR, Korean | korean PP-OCRv5 mobile rec | 13.4 MB | shares the Chinese line detector |
| Erase | Balloon fill, with Telea [23] in a sandbox as fallback | OpenCV.js 12.7 MB | text drawn on art is labelled instead (Section 7.3) |
| Translate | Gemma 2 2B JPN, q4f16 [26] via WebLLM [25], or Gemini Flash-Lite | 1.4 GB (local) | cloud model is opt-in, one request per page |
| Render | Canvas 2D typesetting, lobe-aware layout | – | caption plates with guaranteed contrast |&lt;/p&gt;
&lt;p&gt;Cross-origin isolation (COOP/COEP headers) gives the offscreen document &lt;code&gt;SharedArrayBuffer&lt;/code&gt; and therefore four onnxruntime threads. With them, detection takes about 0.7 s per page and OCR about 0.4 s per region. Single-threaded, detection took 1.9 to 2.7 s, and Japanese OCR took about 1.15 s to encode each bubble plus 1.5 s for every 30 characters decoded. A twelve-bubble page needed roughly 30 s. Telea inpainting takes about 235 ms for a round trip to the sandbox, and layout and drawing take 83 ms together.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;4. Evaluation Methodology&lt;/h2&gt;
&lt;h3&gt;4.1 The reference set&lt;/h3&gt;
&lt;p&gt;We built &lt;code&gt;detect_reference.json&lt;/code&gt; from 57 real pages. Thirty-seven came from pages where earlier sessions had logged failures, plus the OCR reference sets. Twenty were sampled from Danbooru by tag across the three source languages, with a minimum short side of 600 pixels and restricted to general and sensitive ratings. Danbooru shows a licence for each image, so it can be checked at fetch time. The file stores post IDs and boxes rather than pixels, so the set can be rebuilt without redistributing comic pages.&lt;/p&gt;
&lt;p&gt;The ground truth was first proposed by two finders: the baseline YOLO26n at its 0.05 floor, and PP-OCR&apos;s DB text detector. We inverted DB&apos;s unclip in closed form so that its boxes bound the ink rather than the training margin. A reviewer then judged the clustered proposals on contact sheets and added missed regions by hand. Of the original 525 regions, 311 are bubbles, 72 captions and 142 sound effects, and 86 have dark backgrounds. After the corrections in Section 4.4 the final reference has 540 regions: 314 bubbles, 74 captions and 152 sound effects. The reviewer added 58 regions by hand that neither finder had found. One page was left out entirely: it carried dozens of repeated sound effects along its edges and could not be labelled exhaustively. A half-labelled page would have flattered every detector.&lt;/p&gt;
&lt;h3&gt;4.2 Metrics&lt;/h3&gt;
&lt;p&gt;The pipeline crops each detected box for OCR. A box that is too loose costs a little inpainting, while a box that clips the lettering loses words. The primary measure is therefore coverage rather than IoU. For a ground-truth region $g$ and the set of detections $\mathcal{D}$:&lt;/p&gt;
&lt;p&gt;$$
\mathrm{cov}(g) = \frac{\left| g \cap \bigcup_{d \in \mathcal{D}} d \right|}{|g|}, \qquad \text{found}(g) \iff \mathrm{cov}(g) \ge 0.5 .
$$&lt;/p&gt;
&lt;p&gt;A threshold of 0.5 marks the point where OCR starts returning fragments instead of short readings. A detection is &lt;strong&gt;spurious&lt;/strong&gt; when less than 0.25 of its area lies inside any ground-truth region. The cut is generous on purpose, because detector boxes are looser than the lettering inside them. The spurious rate is&lt;/p&gt;
&lt;p&gt;$$
S_m(t) = \frac{#{\text{spurious detections of model } m \text{ at threshold } t}}{#{\text{detections}}} .
$$&lt;/p&gt;
&lt;p&gt;Recall is reported by slice, because the known gaps show up in slices rather than in the mean. The slices are kind (bubble, sound effect, caption), background, and language. A region counts as dark when at least 25% of its pixels fall below luma 80. We also report IoU&gt;0.5 as the conventional localisation measure [41], together with merge and fragmentation counts.&lt;/p&gt;
&lt;h3&gt;4.3 Matched-ceiling comparison&lt;/h3&gt;
&lt;p&gt;Detectors put their confidence on different scales, so a fixed threshold means something different for each model. We sweep thresholds from 0.05 to 0.95 on a 0.01 grid, reusing one inference pass per page. For each spurious ceiling $c$ we then take the threshold that maximises sound-effect recall without exceeding the ceiling:&lt;/p&gt;
&lt;p&gt;$$
t^{*}&lt;em&gt;m(c) = \arg\max&lt;/em&gt;{t ,:, S_m(t) \le c} R^{\text{sfx}}_m(t),
$$&lt;/p&gt;
&lt;p&gt;and we compare models on sound-effect, dark-background and overall recall at $t^{*}_m(c)$. The ceiling that decides shipping is $c = 0.9%$, the baseline&apos;s own rate at its default threshold. Looser ceilings (1.4% to 7%) show how the curves behave where the pipeline would tolerate more errors. The &lt;strong&gt;crossover&lt;/strong&gt; is the lowest ceiling at which a candidate beats the baseline on both sound-effect and dark-background recall.&lt;/p&gt;
&lt;p&gt;Two practical points matter here. First, spurious rate is not monotonic in threshold. Deduplication sees a different set of boxes at each threshold, so a lower threshold can have fewer spurious boxes. Second, the grid has to be fine where the operating points lie. An early grid had a 0.01 step only up to 0.65. Every TextSeg operating point sat above that, and one seed&apos;s apparent failure at 0.9% was an artefact of the coarse grid (Section 6.3).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT]
At a fixed threshold, a detector that emits more boxes looks better on recall. At a matched spurious rate it has to earn its recall with the same false-positive budget. Every shipping decision in this paper uses the matched comparison. Fixed-threshold tables appear only for context.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;4.4 Correcting the pooled reference&lt;/h3&gt;
&lt;p&gt;The reference was proposed by the baseline detector and DB. Anything real that the baseline finds was therefore proposed, reviewed and labelled. Real text that only a new model finds is missing, so it counts against the new model. This is the pooling bias known from information retrieval [27, 28]. We corrected it the only safe way: by hand, one box at a time. The reviewer drew each new box around the ink, not from the candidate detector&apos;s box, and never edited the JSON directly. The added regions were:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;7 regions from reviewing v4&apos;s flagged boxes, 2 from v2&apos;s and 1 from v3&apos;s (525 to 535 regions);&lt;/li&gt;
&lt;li&gt;3 sound effects from TextSeg seed 2&apos;s spurious boxes (538);&lt;/li&gt;
&lt;li&gt;1 sound effect from seed 0&apos;s (539);&lt;/li&gt;
&lt;li&gt;1 sound effect from the recipe 2 gate review (540).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Seven of v4&apos;s fifteen &quot;spurious&quot; boxes turned out to be unlabelled text. On the corrected reference, the genuine spurious rate at threshold 0.25 (535-region reference) is 0.6% for the baseline, 3.3% for v2 and 1.8% for v4. Once a comparison started, we never changed the reference in the middle of it. Every correction was followed by a re-sweep of all models.[^1]&lt;/p&gt;
&lt;h3&gt;4.5 Uncertainty&lt;/h3&gt;
&lt;p&gt;Two sources of noise dominate. &lt;strong&gt;Page sampling&lt;/strong&gt;: the 0.9% gate rests on two or three boxes out of about 350. With about 150 sound-effect regions (147 at the 535-region stage, 152 in the final reference), the 95% interval on sound-effect recall is about ±0.075, which is as large as most differences between models. We therefore report a page bootstrap [30]: 1,000 resamples of the 57 pages, with the threshold re-chosen inside each resample, giving $P(\text{candidate beats the baseline on sfx and dark})$, where beating means strictly higher recall on both. &lt;strong&gt;Training seeds&lt;/strong&gt;: we report every seed separately and compare distributions, never single runs [29, 44].&lt;/p&gt;
&lt;h3&gt;4.6 Leakage guards&lt;/h3&gt;
&lt;p&gt;The training, negative-mining and expansion sets are all sampled with the reference&apos;s post IDs excluded (53 IDs were refused on the first training draw). The dataset builder and the training notebook both re-assert that no reference page is present. We also tested the guards by planting a reference page into a synthetic dataset directory and confirming that both build scripts refused it. When a review found informative false positives on reference pages, those boxes did not become training negatives. Only their categories informed the next mining round on fresh pages.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;5. Adapting an Off-the-Shelf Detector: Six YOLO Runs&lt;/h2&gt;
&lt;h3&gt;5.1 Baseline and the two gaps&lt;/h3&gt;
&lt;p&gt;At the 0.05 floor, the baseline found 0.974 of bubbles but only 0.415 of sound effects and 0.477 of dark-background regions. At its default threshold of 0.25 those two figures fell to 0.155 and 0.384. Both gaps are ones readers notice: sound effects are left untranslated, and dialogue on night scenes or dark panels is skipped. Both gaps are also missing from a mean score, because bubbles make up most of the reference.&lt;/p&gt;
&lt;h3&gt;5.2 A training set built without exhaustive labelling&lt;/h3&gt;
&lt;p&gt;We sampled 2,925 Danbooru pages from outside the reference set, weighted toward the gaps. The weights were 1,170 sound-effect pages, 720 dark, 300 screentone, 540 Korean or Chinese, and 270 plain monochrome pages as a control, so that sound-effect recall could not be bought by forgetting bubbles. Hand-labelling at that scale was not possible, so labels came from three independent finders. The finders were proposers, and none of them was trusted as ground truth.&lt;/p&gt;
&lt;p&gt;Two finders were not enough. On the original reference, the baseline and DB together proposed 90% of regions, but only 20% of the regions the reviewer had added by hand. Measured across all sound effects, 43 of 142 were found by neither. That mattered, because an unlabelled region is a negative in YOLO&apos;s loss, so a set built from those two finders would have taught the model to suppress the very sound effects we wanted it to find. Adding ogkalu&apos;s RT-DETR comic detector [47] as a third finder closed most of the hole. Its class for text outside balloons covers exactly the missing shape.&lt;/p&gt;
&lt;p&gt;Clusters of proposals were then kept or dropped by a rule, &lt;code&gt;autolabel.py&lt;/code&gt;, fitted against a hand-reviewed tier: 30 pages at first, 42 later, and 103 after the expansion described below. The first rule used finder confidence and box size. It labelled at precision 0.920 and recall 0.882 overall, and 0.766 / 0.735 on sound effects.&lt;/p&gt;
&lt;h3&gt;5.3 v1: recall solved, precision created&lt;/h3&gt;
&lt;p&gt;Fine-tuning &lt;code&gt;Kiuyha/Manga-Bubble-YOLO&lt;/code&gt; with Ultralytics [45] on this set took 20 epochs on a free Colab T4 (imgsz 1280, batch 8, AdamW, learning rate 0.002, 1.06 h). Two engineering fixes came first. Pages as large as 79 megapixels made data loading take 96 s per iteration until the export capped the long side at 1,600 pixels. Unthrottled fetching triggered HTTP 429 after about 400 pages and silently shrank the dataset until we added a delay and backoff. We also confirmed that the export path changed nothing: re-exporting the base checkpoint reproduced the baseline&apos;s scores exactly, so any later change is due to training.&lt;/p&gt;
&lt;p&gt;At threshold 0.25, v1 raised sound-effect recall from 0.155 to 0.465 and dark-background recall from 0.384 to 0.558. The spurious rate rose from 0.9% to 7.8%. The threshold sweep in Table 2 shows what that means.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Table 2. Fine-tune v1 against the baseline on the 525-region reference. Each row compares the two models at a similar spurious rate.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| Spurious rate | Baseline (threshold → sfx / dark / all) | v1 (threshold → sfx / dark / all) |
|---|---|---|
| ≈4.2% | 0.05 → 0.415 / 0.477 / 0.792 | 0.40 → 0.366 / 0.465 / 0.701 |
| ≈3.4% | 0.10 → 0.296 / 0.453 / 0.741 | 0.425 → 0.345 / 0.430 / 0.688 |
| ≈1.9% | 0.15 → 0.218 / 0.419 / 0.705 | 0.475 → 0.275 / 0.349 / 0.640 |
| ≈0.8% | 0.25 → 0.155 / 0.384 / 0.659 | 0.65 → 0.113 / 0.244 / 0.512 |&lt;/p&gt;
&lt;p&gt;At equal false-positive cost, v1 loses on overall recall, bubble recall and IoU at every point. It wins on sound effects only in a narrow band between 2% and 3.4%, and by about 0.05 rather than the 0.31 the fixed-threshold view suggested. The largest improvement available without training was simply to lower the baseline&apos;s own threshold. The baseline at 0.05, which users could already select, beat v1 at 0.40 on every slice at the same spurious rate. The model had learned the label noise as confident text. Because the noise did not sit at low confidence, no threshold could remove it.&lt;/p&gt;
&lt;h3&gt;5.4 v2: the label rule was reading the wrong variable&lt;/h3&gt;
&lt;p&gt;We broke the rule&apos;s false labels down by which finders had proposed each cluster, instead of tightening its thresholds. On the reviewed pages the base rates were far apart:&lt;/p&gt;
&lt;p&gt;| Finders that proposed the cluster | Real text |
|---|---|
| Baseline YOLO, with or without others | 326 / 334 = 0.976 |
| DB and RT-DETR only | 36 / 52 = 0.692 |
| DB only | 10 / 28 = 0.357 |
| RT-DETR only | 25 / 418 = 0.060 |&lt;/p&gt;
&lt;p&gt;Fourteen of the old rule&apos;s 22 false labels came from lone RT-DETR proposals. Inside that group RT-DETR&apos;s confidence was inverted: the clusters the old rule kept scored 0.16 to 0.36 and were mostly art, while the real ones it dropped scored 0.05 to 0.14 and were large sound effects. DB proposals, for their part, arrive with no score at all, so the old score cut had silently thrown away every cluster DB found alone. The rebuilt rule looks at corroboration first and size second. Every YOLO-backed cluster is kept, DB-and-RT-DETR clusters need 2,500 px², DB-only clusters need 1,200 px², and anything else needs a lone finder at confidence 0.50 or higher. False labels fell from 22 to 6 while true positives rose from 254 to 256. A page-level bootstrap gave a precision gain of +0.057 (95% CI [+0.027, +0.088]). In split-half cross-validation, where the constants were fitted on 15 pages and scored on the other 15, the gain was +0.048 and positive in 99% of splits.&lt;/p&gt;
&lt;p&gt;The reviewed set had included no screentone pages, and those were where the new rule changed the most. Twelve screentone pages were reviewed to cover them. The old rule&apos;s precision on them was 0.849, against 0.920 elsewhere, and the new rule raised it to 0.925. Over all 42 reviewed pages, precision went from 0.899 to 0.961 and sound-effect precision from 0.695 to 0.878.&lt;/p&gt;
&lt;p&gt;The relabelled set had 2,926 pages and 23,480 regions. We retrained v2 with the same budget as v1 (1.055 h against 1.056 h), so only the labels differed between the two runs. At threshold 0.25 the spurious rate fell from 7.8% to 4.6%, and every recall slice rose. That combination is what a label-quality fix looks like, as opposed to trading recall for precision. At the 0.9% ceiling, however, v2 still lost: it reached sound-effect 0.085 and dark 0.233 at threshold 0.65, against the baseline&apos;s 0.155 and 0.384. Its crossover sat near 2.5%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fragmentation was a second symptom, not a second problem.&lt;/strong&gt; The fine-tunes looked more fragmented: 13 fragmented regions against the baseline&apos;s 4. The metric counted a region as fragmented whenever two detections each covered more than 5% of it, which also fires when a neighbouring box clips its edge. Only 5 of v2&apos;s 13 were genuine splits. When the three models were matched on detection count rather than threshold (346, 346 and 344 boxes), the ordering reversed to 4, 3 and 2. The end-to-end head was intact, since duplicates before deduplication stayed at 0.047 to 0.058 per box, and removing the parent-drop rule changed nothing. One separable effect did appear: annotation granularity. The fine-tunes drew one box per vertical column where the baseline drew one per balloon, and each model reproduced the column-pair rate of its own training labels.&lt;/p&gt;
&lt;h3&gt;5.5 v3: the starting checkpoint is not the cause&lt;/h3&gt;
&lt;p&gt;v1 and v2 both started from the manga-tuned checkpoint the extension already shipped. To test whether that starting point was holding the operating point back, and to move toward a model the project owned end to end, we retrained from Ultralytics&apos; generic COCO checkpoint on the same data. Two details show how easily a run&apos;s meaning can shift. First, &lt;code&gt;optimizer=&apos;auto&apos;&lt;/code&gt; would silently have switched from AdamW to SGD at 40 epochs, so we pinned the optimizer. Second, Ultralytics 8.4 changed the export default, so the NMS-free head is exported only when &lt;code&gt;nms=False&lt;/code&gt; is passed. Our shape assertion caught the resulting &lt;code&gt;[1, 5, 33600]&lt;/code&gt; output before any scoring ran.&lt;/p&gt;
&lt;p&gt;v3 needed about 27 epochs to reach the validation fitness v2 reached in 20. It ended almost on top of v2: 5.1% spurious at 0.25, with the same shape of curve. Two very different starting points gave one answer, so the confident false positives come from the data and labels, not from the initial weights.&lt;/p&gt;
&lt;h3&gt;5.6 v4: hard negatives are the first lever that moves the curve&lt;/h3&gt;
&lt;p&gt;A later audit split the remaining label errors by page type. Screentone and hatching accounted for 7 of 15 false labels on the reviewed pages, marks backed by YOLO for 4, fragments for 2, and painterly colour pages for 2. We tested the hypothesis that dense, painterly art causes the false positives. Colourfulness [43] did not support it: precision was 0.940 on greyscale pages and 0.972 on colour pages.&lt;/p&gt;
&lt;p&gt;So we mined the models&apos; own confusions. We ran the baseline, v2 and v3 over 700 fresh pages, weighted toward textures, effect marks and rendered art, and kept every box at threshold 0.25. That gave 7,795 candidates. The two text finders were used only to order the review. Tier A (no finder agrees) turned out to be 64% negatives, tier B (only RT-DETR agrees) 27%, and tier C (DB agrees) 1%. A and B were reviewed exhaustively and C was sampled. Every kept crop was then checked again at full size, because two crops that looked clean on the contact sheet turned out to hold clipped lettering.&lt;/p&gt;
&lt;p&gt;The result was 158 confirmed negative crops on 123 pages. Of these, 94 were effect and costume marks, 63 rendered art, 1 a fragment, and 0 screentone. Two of the diagnosed categories could not be reached this way. A fragment always has lettering inside any honest window around it. Screentone never appeared as a confident false positive on fresh pages. The crops were packed onto grey canvases as background images, making up 5% of the training set, and v2&apos;s recipe was otherwise left unchanged.&lt;/p&gt;
&lt;p&gt;v4 was the first run to change the shape of the curve. At threshold 0.25 its spurious rate was 3.3% and every recall slice rose over v2. Genuine false positives at 0.25 fell from 15 to 8, and the ones that disappeared were exactly the marks and art the negatives targeted. Seven of v4&apos;s remaining fifteen &quot;spurious&quot; boxes were unlabelled text (Section 4.4).&lt;/p&gt;
&lt;h3&gt;5.7 v5 and v4b: the seed moves results as much as the recipe&lt;/h3&gt;
&lt;p&gt;A second mining round against v4 on 200 fresh pages found only 51 negatives. Most of what v4 now flagged was real text: in the DB-backed tier, 92 of 95 candidates were. The easy negatives were used up. We therefore expanded the true positives instead. Sixty-one fresh pages were reviewed with v4 as a fourth proposer: 587 regions were kept, 652 candidates rejected and 54 added by hand. The result was v5. v5 lost to v4 at every ceiling, and two cheap explanations failed. All 11 of v5&apos;s spurious boxes were genuine errors. The expansion regions that only v4 had proposed improved the least, which is the opposite of what circularity would predict.&lt;/p&gt;
&lt;p&gt;We then reran v5&apos;s exact recipe with seed 1337, since v4 and v5 had both used seed 0 according to their published &lt;code&gt;args.yaml&lt;/code&gt;. This replicate, v4b, settled the question (Table 3).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Table 3. Matched-ceiling recall (sfx / dark) on the 535-region reference. v5 and v4b are the same recipe with different seeds.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| Ceiling | Baseline | v2 | v3 | v4 | v5 | v4b |
|---|---|---|---|---|---|---|
| 0.9% | 0.150 / 0.384 | 0.136 / 0.279 | 0.265 / 0.384 | 0.265 / 0.372 | 0.184 / 0.326 | &lt;strong&gt;0.327 / 0.442&lt;/strong&gt; |
| 1.9% | 0.259 / 0.442 | 0.293 / 0.395 | 0.333 / 0.442 | &lt;strong&gt;0.544 / 0.663&lt;/strong&gt; | 0.463 / 0.581 | 0.381 / 0.523 |
| 2.8% | 0.265 / 0.442 | 0.449 / 0.547 | 0.435 / 0.547 | &lt;strong&gt;0.612 / 0.709&lt;/strong&gt; | 0.524 / 0.674 | 0.503 / 0.628 |
| Crossover | – | 1.39% | – | 0.96% | 1.32% | 0.81% |&lt;/p&gt;
&lt;p&gt;The two runs of the identical recipe disagree by as much as v4 and v5 do: the same-recipe spread in sound-effect recall is 0.143 at 0.9% and 0.082 at 1.9%. The order also changes with the ceiling. v4b is the best of the three at 0.9% and the worst at 1.9% and 2.8%. No single 20-epoch run is evidence about its recipe. We did not promote v4b, because doing so would repeat the mistake the replicate had just exposed. The baseline stayed in the extension, since the pipeline&apos;s 0.9% tolerance sits inside this noise band.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;gitGraph
    commit id: &quot;baseline yolo26n&quot;
    branch manga-yolo-lineage
    commit id: &quot;v1 noisy labels&quot;
    commit id: &quot;v2 corroboration relabel&quot;
    commit id: &quot;v4 hard negatives&quot;
    commit id: &quot;v5 plus 61 pages&quot;
    commit id: &quot;v4b seed 1337&quot;
    checkout main
    branch coco-start
    commit id: &quot;v3 COCO checkpoint&quot;
    checkout main
    branch textseg
    commit id: &quot;R1 three seeds&quot;
    commit id: &quot;R2 fixed val, EMA&quot;
    commit id: &quot;R3 own negatives&quot;
    checkout main
    merge textseg id: &quot;ship TextSeg&quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Figure 2. Detector lineage. Each commit is a trained and evaluated condition, and the three TextSeg commits are three seeds each.&lt;/em&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!NOTE]
&lt;strong&gt;What the YOLO lineage established.&lt;/strong&gt; The labels were a cause, and the corroboration rule fixed part of it. The starting checkpoint was not a cause. Hard negatives were the first lever to move the curve. At this data and compute scale, run-to-run variance was as large as the differences between recipes, so we compared distributions rather than single runs from then on.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;6. TextSeg: A Text-Probability-Map Detector Trained With an Ignore Mask&lt;/h2&gt;
&lt;h3&gt;6.1 Why write a detector at all&lt;/h3&gt;
&lt;p&gt;By this point the YOLO lineage had run into four limits. &lt;strong&gt;(i)&lt;/strong&gt; The shipping test could not tell runs apart: at 0.9% on the 535-region reference, the baseline scored 2 false boxes of 347, v4 3 of 349, and v4b 3 of 372, so one box decided the outcome. &lt;strong&gt;(ii)&lt;/strong&gt; 2,884 of the 2,987 training pages (96.6%) carried automatic labels, and the labeller finds only about 70% of sound effects. The rest were being taught as background, and Ultralytics offers no way to leave a region out of the loss. &lt;strong&gt;(iii)&lt;/strong&gt; Every fine-tune scored about 0.61 on IoU&gt;0.5 against 0.88 for the baseline, even as coverage rose, because the training boxes come from finders that draw a line, a text block or a whole balloon for the same lettering. &lt;strong&gt;(iv)&lt;/strong&gt; Ultralytics states that its AGPL-3.0 licence covers models its training code produces, so every weight in the lineage inherited that provenance.&lt;/p&gt;
&lt;p&gt;A per-pixel text map deals with (ii) to (iv) together. A column box and a block box paint mostly the same pixels. An ignore mask becomes one multiplication in a loss we write ourselves. The model, the loss and the training loop are this project&apos;s own code, on an Apache-2.0 ImageNet backbone.&lt;/p&gt;
&lt;h3&gt;6.2 Model, labels and loss&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; A timm MobileNetV3-Large backbone [11] feeds features at strides 4, 8, 16 and 32 into an FPN [12] with 96 channels. Each level is smoothed to 24 channels, upsampled to stride 4, concatenated, and passed to a head that outputs one logit map. The head&apos;s bias starts at $-4$ ($\sigma(-4) \approx 0.02$), because nearly every pixel is background. Normalisation is built into the graph, so the ONNX file takes the same 0-to-1 input tensor, &lt;code&gt;[1, 3, 1280, 1280]&lt;/code&gt;, as the YOLO export did, and outputs &lt;code&gt;[1, 1, 320, 320]&lt;/code&gt;. The file is 12.8 MB, and ONNX export drifts from PyTorch by less than $10^{-5}$.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labels.&lt;/strong&gt; Each text box is painted shrunk in the DBNet manner [9], with every side moved in by&lt;/p&gt;
&lt;p&gt;$$
D = \frac{A,(1 - r^2)}{L}, \qquad r = 0.4,
$$&lt;/p&gt;
&lt;p&gt;where $A$ is the box area and $L$ its perimeter, so balloons that touch stay separate blobs. Every pixel of a training crop is one of four kinds. It is text if it lies inside a shrunk core. It is background if it lies elsewhere on the page. It is ignored if it lies inside a region the automatic labeller dropped even though DB had proposed it: such clusters are real text 36% to 69% of the time, so neither answer is safe to teach. Lone RT-DETR proposals and thin slivers are real only 6% and 1 in 29 of the time, so they stay background. The fourth kind is the ring between each core and its full box. At first we left the ring out of the loss. The smoke test showed that this let the model paint whole boxes, which the decoder then grew a second time, so the ring is now supervised as background, as DBNet does it. The full set holds 24,067 text regions and 1,848 ignore regions. Hard-negative pages enter whole, with everything except the confirmed windows ignored.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Loss.&lt;/strong&gt; The loss is binary cross-entropy over masked pixels, keeping at most three hardest negatives per positive [13], plus a masked dice term [14]:&lt;/p&gt;
&lt;p&gt;$$
\mathcal{L} = \frac{\sum_{p \in P} \ell_p + \sum_{n \in \mathrm{top}_{3|P|}(N)} \ell_n}{|P| + \min(|N|, 3|P|)} ;+; 1 - \frac{2\sum m,\hat{y},y + 1}{\sum m,\hat{y} + \sum m,y + 1},
$$&lt;/p&gt;
&lt;p&gt;with $m$ the loss mask and $|P|$ floored at 1, so that a crop with no text still pushes its background down. &lt;strong&gt;Decoding.&lt;/strong&gt; The map is thresholded at 0.3, and each 4-connected component of at least 3 cells becomes a region, scored by its mean probability. Its box is grown back by the unshrink offset, which is found by solving the same fixed-point equation as training. A TypeScript port of the decoder in the extension gave identical boxes to the Python version on all 57 reference maps (1,136 boxes). The parity check fails if a single constant is changed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Training.&lt;/strong&gt; We trained on Kaggle [42] because a Kaggle run continues after the browser closes, so no machine has to stay on during training. The GPU was a T4 ×2, with AdamW at learning rate $10^{-3}$, weight decay $10^{-4}$, 768-pixel crops, batch 8, one warm-up epoch and cosine decay to 2%. Pages are cached once at a 1,536-pixel long side (2,987 pages plus 160 negative pages, 1.2 GB, no failed downloads), so each run reads from the cache instead of re-downloading. A CPU smoke test first checked that the model learns at all: after 90 s of training on 10 pages, pixel F1 on 2 unseen reviewed pages rose from 0.64 to 0.72.&lt;/p&gt;
&lt;h3&gt;6.3 Three recipes, nine runs&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Table 4. TextSeg matched-ceiling recall on the 540-region reference: range over seeds, with the mean. Every row, including the bootstrap row, is computed on the 540-region reference (1,000 page resamples; &quot;beats&quot; means strictly higher on both sound effects and dark backgrounds).&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| Ceiling, slice | Baseline | R1: 80 ep, best-val epoch | R2: last-epoch EMA | R2: val-peak EMA | R3: R2 + own negatives |
|---|---|---|---|---|---|
| 0.9% sfx | 0.145 | 0.289–0.461 (0.390) | 0.289–0.414 (0.340) | 0.270–0.336 (0.307) | &lt;strong&gt;0.375–0.441 (0.408)&lt;/strong&gt; |
| 0.9% dark | 0.379 | 0.414–0.598 (0.525) | 0.425–0.517 (0.460) | 0.414–0.471 (0.441) | &lt;strong&gt;0.506–0.563 (0.533)&lt;/strong&gt; |
| 1.9% sfx | 0.250 | 0.487–0.553 (0.518) | 0.434–0.553 (0.480) | &lt;strong&gt;0.572–0.592 (0.586)&lt;/strong&gt; | 0.539–0.592 (0.568) |
| 1.9% dark | 0.437 | 0.563–0.678 (0.613) | 0.529–0.598 (0.563) | &lt;strong&gt;0.655–0.678 (0.670)&lt;/strong&gt; | 0.655–0.667 (0.663) |
| 4.3% sfx | 0.395 | 0.539–0.691 (0.629) | 0.605–0.658 (0.627) | 0.658–0.743 (0.700) | &lt;strong&gt;0.717–0.743 (0.730)&lt;/strong&gt; |
| 4.3% dark | 0.471 | 0.632–0.793 (0.686) | 0.655–0.724 (0.693) | 0.713–0.793 (0.751) | &lt;strong&gt;0.782–0.805 (0.793)&lt;/strong&gt; |
| $P$(beats baseline at 0.9%), seeds 0/1/2 | – | 0.65 / 0.97 / 1.00 | – | 0.82 / 0.71 / 0.69 | &lt;strong&gt;0.97 / 0.92 / 0.97&lt;/strong&gt; |&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Recipe 1&lt;/strong&gt; used seeds 0, 1 and 2 for 80 epochs each, at about 82 s per epoch and 1 h 53 min to 2 h per run, and exported the epoch with the best validation F1. On the 535-region reference, as first reported, seeds 1 and 2 cleared the 0.9% gate and seed 0 did not. An independent audit of the raw logs overturned three readings of that result:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;em&gt;The seeds were validated on different pages.&lt;/em&gt; The 25 validation pages were drawn with the training seed, so any two seeds shared only 4 to 8 of them. Validation F1 differed by 0.05 before training had even diverged. The apparent ordering in which &quot;val F1 predicts the gate&quot; (0.7085 &amp;#x3C; 0.7367 &amp;#x3C; 0.7544) could not be compared across seeds, and with three points a monotone order happens by chance one time in three.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Seed 0&apos;s failure was partly a grid artefact.&lt;/em&gt; Re-swept at 0.01 steps up to 0.95 on the 535-region reference, seed 0 does reach the 0.9% ceiling (threshold 0.81, 2 of 292 spurious). Its reported crossover of 1.14% could not be reproduced (0.97%).&lt;/li&gt;
&lt;li&gt;&lt;em&gt;What set seed 0 apart was unsupervised padding.&lt;/em&gt; The loss masked out the grey letterbox, so the model&apos;s output there was never trained. A component straddling the page edge survived decoding as a full-height sliver at $x = 0$. These slivers were 46 of seed 0&apos;s 83 spurious boxes at 0.25, against 15 for seed 1, 10 for seed 2 and none for the baseline. An unconstrained region is exactly where seeds should diverge.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The same audit tested a hypothesis that had been recorded as fact: that the training boxes were drawn larger than the reference boxes. The gap in size relative to the page is real. The median Japanese bubble box is 10.1 to 10.8 per mille of the page in training against 6.0 in the reference. But it reflects the pages, not the labelling. Boxes drawn by the same baseline detector at the same floor are 10.3 per mille on training pages and 5.9 on reference pages, because the training pages simply have larger balloons (median box side 194 px against 134 px). Most of TextSeg&apos;s low IoU instead came from its own decoder (Section 6.5). The ignore mask, for its part, has no clean evidence either way. Recall above the labeller&apos;s own sound-effect recall is not diagnostic, since v4 exceeds it too without any mask. An ablation would need many seeds per arm while the seed spread stays near 0.14.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Recipe 2&lt;/strong&gt; changed four things at once. It supervised the padding as background, fixed the validation pages across runs, cut training to 40 epochs (the best epochs had been 27 to 41), and kept an exponential moving average of the weights [31] with decay 0.999. All three seeds ran in a single 3 h 13 min Kaggle session. The page-edge slivers disappeared: zero at every threshold, against 46, 15 and 10 in recipe 1. The EMA&apos;s validation F1 peaked at epoch 22 or 23 in every seed, and the last-epoch export had already lost recall. Exporting the EMA at its validation peak gave the tightest condition yet from 1.9% upward, with a seed spread of 0.020 in sound-effect recall against 0.066 for recipe 1. At 0.9%, however, nothing could be resolved: bootstrap probabilities were only 0.69 to 0.82.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The gate was the same few boxes in every seed.&lt;/strong&gt; Across the three seeds, the spurious boxes at the 0.9% and 1.9% operating points came to 15 distinct boxes, and all were reviewed by hand. Thirteen were genuine errors: 8 were art, such as a &quot;100&quot; badge, a doodled face, a chibi and a shirt logo, and 5 were marks, such as scribbles, a ♪ and a running-man icon. One was text not meant for translation, and one was a real sound effect missing from the reference. All six boxes that set the 0.9% ceilings were genuine, and four of them fired in all three seeds. Three of the six sat on the reference set&apos;s sketch pages, which carry only one or two regions each. The gate was therefore neither seed noise nor missing labels. It was a handful of glyph-like shapes that every run ranked above real text.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Recipe 3&lt;/strong&gt; targeted those shapes on fresh pages, never on the reference pages themselves. A Kaggle CPU notebook ran the recipe 2 models over 500 new pages and excluded 4,282 post IDs already used anywhere. The pages were weighted toward doodles, badges, scribbles, symbol marks and hatching. Punctuation-only bubbles were excluded from the weighting, because the reference counts &quot;?&quot; and &quot;!&quot; as text. Candidates scoring 0.5 or more went onto 17 review sheets (510 crops). The negatives were then checked again at full resolution, and this round kept 200 of them. Retraining with all three negative lists, from 3,263 training records including 301 negative pages, moved the 0.9% point on every seed. Each recipe 3 seed beats every recipe 2 seed on sound-effect and dark-background recall at the gate. The operating threshold fell from 0.74–0.76 to 0.70. The bootstrap probability of beating the baseline rose to 0.97, 0.92 and 0.97 for seeds 0, 1 and 2. The six gate boxes still fire, but at lower scores: the top one fell from 0.85 to 0.72.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!WARNING]
&lt;strong&gt;What this does not show.&lt;/strong&gt; For seeds 0 and 2 the 95% interval of the dark-background improvement at 0.9% touches zero, and for seed 1 it crosses zero. The gate still rests on two or three boxes out of about 390. We do not rank TextSeg against v4 or v4b: they differ in architecture, in labels and in seed variance, and no controlled comparison separates those factors.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;6.4 Seed variance, measured directly&lt;/h3&gt;
&lt;p&gt;Two recipes let us compare within-recipe spread to between-recipe differences. At 0.9%, recipe 1&apos;s seeds spanned 0.29 to 0.46 in sound-effect recall. That range is wider than any difference between recipes measured in this project. Fixing the validation split and exporting a smooth EMA narrowed the spread above 1.9% threefold. Only the targeted negatives moved the 0.9% point on every seed. The practical rule we took from this: a recipe gets judged on three seeds and on page-bootstrap intervals, never on one run and never at a fixed threshold. Seed 2 of recipe 1 illustrates why. Its sound-effect recall at 0.25 was exactly 109/147, the same as seed 0&apos;s, yet at matched ceilings they were the best and the worst of the three.&lt;/p&gt;
&lt;h3&gt;6.5 Box extent and the decoder&lt;/h3&gt;
&lt;p&gt;The score threshold only filters whole components, so box extent is set entirely by the binarisation cut and the unshrink offset. That is why IoU stays flat as the threshold changes. We ran a decoder-only grid over &lt;code&gt;BINARIZE&lt;/code&gt; ∈ {0.2, 0.3, 0.4, 0.5} and unshrink scales from 1.25 down to 0.25, on all three recipe 3 seeds, with no retraining. The unshrink scale traded coverage against IoU and nothing else. IoU&gt;0.5 rose from 0.36 to 0.53 and then 0.67 as the scale fell from 1.0 to 0.75 and then 0.5, and every step cost coverage and matched-ceiling recall. The oversized boxes had been compensating for cores that sit poorly on the lettering. Since coverage is the primary metric, the decoder stayed at 0.3 with the full offset. Setting 0.2 with ×0.75 kept coverage (0.909 against 0.906) and added 0.13 IoU, so we noted it for erasing, where box extent matters.&lt;/p&gt;
&lt;h3&gt;6.6 Erasing with the map, not the box&lt;/h3&gt;
&lt;p&gt;In the browser TextSeg found more text than YOLO, including ぬいっ, ガラ, スー, ビギッ and handwritten asides, but it erased worse. The erase stage relied on each box sitting inside its balloon so that Telea fills from clean interior, and TextSeg&apos;s looser boxes crossed balloon outlines, which Telea smeared inward. The raw map could not replace the box, because it paints a shrunk core per text block rather than individual strokes. Erasing only the map&apos;s cells left 13% to 36% of the lettering behind at every threshold and dilation we tried. Instead we grow each component&apos;s core by a disk of 0.75 times its own unshrink offset. This rebuilds the box minus its corners, and the corners are where a rectangle crosses a round balloon. Section 7.3 gives the measurements.&lt;/p&gt;
&lt;p&gt;TextSeg ships as the extension&apos;s only detector at a default threshold of 0.60. This is looser than the 0.9% gate (threshold 0.70 for the deployed seed). We picked it by inspecting real pages in the development harness, where 0.5 to 0.6 added real lettering and almost no false boxes. It also sits near the model&apos;s matched 1.9% point on the reference (0.61). The default was not derived from the gate, and calibrating it per language against the matched sweep is listed as open work (Section 8.3).&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;7. Downstream Stages: Decisions Fixed by Measurement&lt;/h2&gt;
&lt;p&gt;The detector sets the operating point, and the downstream stages set what each region costs. This section reports the measurements behind each stage&apos;s design. We give them briefly, because each one follows the same pattern: a plausible default, then a check on real data that confirmed it or overturned it.&lt;/p&gt;
&lt;h3&gt;7.1 Recognition: the script chooses the engine&lt;/h3&gt;
&lt;p&gt;We sent Chinese and Korean text through manga-ocr to see how it would fail. It did not degrade gracefully: the output was unrelated to the input. Its vocabulary lacks 这, 种 and 门, and it contains only 191 of the 11,172 Hangul syllable blocks. Worse, it is confident when wrong: it assigned mean log-probabilities of −0.04 to −0.78 to garbled Chinese, mostly above the −0.5 cut our harness used to flag weak readings. That is why the source language is a user setting and not an automatic guess. A wrong guess produces fluent nonsense and gives no sign that it failed.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Table 5. OCR by script, measured on regions the detector found on real pages.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| Script, set | Engine | Exact | CER |
|---|---|---|---|
| Japanese, 12 reference crops [46] | manga-ocr, int8/int8 (111 MB) | 10/12 | 0.041 |
| Japanese, same, without post-processing | manga-ocr | 6/12 | – |
| Japanese, 24 multi-column balloons | whole region | 16/24 | 0.047 |
| | chunks of ≤ 18 characters (released setting) | &lt;strong&gt;21/24&lt;/strong&gt; | &lt;strong&gt;0.004&lt;/strong&gt; |
| Chinese, 16 horizontal regions | manga-ocr | 1/16 | 0.420 |
| | PP-OCRv6 small det + small rec (31 MB) | &lt;strong&gt;13/16&lt;/strong&gt; | &lt;strong&gt;0.027&lt;/strong&gt; |
| Chinese, 6 vertical columns | PP-OCR, fixed-pitch cells | 0/6 | 0.333 |
| | PP-OCR, ink-projection cells (Otsu [32]) | 2/6 | 0.253 |
| Korean, 67 regions on 11 pages | manga-ocr | 0/67 | 1.080 |
| | PP-OCRv6 Chinese recogniser | 0/67 | 0.937 |
| | korean PP-OCRv5 mobile (13.4 MB) | &lt;strong&gt;21/67&lt;/strong&gt; | &lt;strong&gt;0.153&lt;/strong&gt; |&lt;/p&gt;
&lt;p&gt;Several details in this table came from measurement rather than design. Porting manga-ocr&apos;s post-processing, which folds half-width to full-width characters, doubled the number of exact readings. For Chinese, the choice of detector mattered more than the size of the recogniser: swapping in the older v5 line detector cost more than dropping from the medium recogniser to the small one. For Korean, two script-specific fixes mattered. Joining lines with a space lifted exact readings from 10 to 21. Line grouping had to compare box centres rather than overlap, because DB&apos;s unclip grows every line box by about 22 px, so neighbouring lines always overlap. Vertical Korean never appeared: all 83 Korean regions on 11 pages were horizontal. We therefore report vertical Korean as untested rather than handled. For Japanese, manga-ocr squashes each crop to 224×224 pixels, so a four-column balloon leaves each character only a few pixels. Reading long regions in groups of at most 18 characters, using PP-OCR&apos;s line detector to find the columns, fixed most multi-column misreads without changing the 12 standard crops.&lt;/p&gt;
&lt;h3&gt;7.2 Translation: the fallback that makes it reachable&lt;/h3&gt;
&lt;p&gt;The on-device model was chosen from WebLLM&apos;s 163 prebuilt models on the same 13 real OCR lines. On that set Qwen2.5-1.5B made a tenfold error in an amount of money and left a kanji untranslated in its English output. Gemma 2 2B JPN made neither error, at 1.49 GB against 0.88 GB. WebGPU is not universal, so we added a cloud fallback. We compared four providers on the constraints that decide the choice. CORS turned out not to matter, since the pipeline document is exempt anyway. The deciding factor was the free tier: only Gemini had one, and a fallback that needs a credit card would not help the users who lack WebGPU. DeepL also cannot be prompted, so it cannot follow the instruction to render sound effects as English sound effects.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Table 6. Meaning errors on 23 real OCR lines (12 Japanese, 5 Chinese, 6 Korean), same prompt and clean-up for both models.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| | Japanese | Chinese | Korean | of which SFX | All | ms/line |
|---|---|---|---|---|---|---|
| gemma-2-2b-jpn-it (local, WebGPU) | 6 | 3 | 4 | 2 of 2 | 12 | 2946* |
| gemini-3.5-flash-lite (cloud) | 0 | 1 | 0 | 0 of 2 | 1 | 84 |&lt;/p&gt;
&lt;p&gt;*Measured on a laptop GPU that was held at 210 MHz of its 2,100 MHz by a power policy (0.6 TFLOPS against an expected 15 to 18). The latency figure describes that machine, not the model.&lt;/p&gt;
&lt;p&gt;The local model changed a place name, dropped a chapter heading, turned a scream and a doorbell into unrelated exclamations, and reversed grammatical person in Korean. These are errors of content, not the tone-and-idiom limitations the project had expected when it started. Korean is the weakest language, as expected for a Japanese-tuned model. The cloud path required one more change to be usable at all. Sending one request per bubble got 7 of 23 lines rejected with HTTP 429 on the free tier. Sending one numbered request per page got none rejected and was 17 times faster per line. It was also slightly better, because the model saw the whole page&apos;s dialogue. Three more practical facts about the cloud API: &lt;code&gt;temperature: 0&lt;/code&gt; did not make its output reproducible, a 96-token budget produced empty replies from models that spend tokens on reasoning, and the models-list endpoint was not a reliable sign that a model was available. When a selected target language is not supported by the local model, the extension stops and offers cloud translation. It never switches silently, because that would send the user&apos;s text to a third party they had not chosen.&lt;/p&gt;
&lt;h3&gt;7.3 Erasing: follow the balloon&lt;/h3&gt;
&lt;p&gt;We measured the erase stage on the reference pages with two quantities. &lt;em&gt;Letters erased&lt;/em&gt; is the share of dark pixels (luma below 128) belonging to lettering that stays inside its reference box, and it should be high. &lt;em&gt;Other ink&lt;/em&gt; is the number of dark pixels erased outside the lettering per page, which is damage to outlines and art, and it should be low.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Table 7. Erase masks on the 57 reference pages.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;| Mask | Letters erased | Other ink erased (px/page) |
|---|---|---|
| YOLO box, padding 0 | 0.985 | 2,656 |
| TextSeg box | 0.963 | 36,894 |
| TextSeg shape, grow ×0.75 | 0.967 | 12,403 |
| YOLO box + balloon fill | 0.985 | 892 |
| TextSeg shape + balloon fill | 0.984 | 9,038 |&lt;/p&gt;
&lt;p&gt;Box padding was a case where the obvious choice was wrong. We started with 3 px of padding for anti-aliasing halos. Measured residual ink rose steadily with padding: 7.3% at 0 px, 19.3% at 3 px and 38.4% at 8 px. Padding pushes the ring Telea samples from off the clean balloon interior and onto the dark outline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balloon fill&lt;/strong&gt; handles a case from a benchmark page used since the start of the project: a single line of speech running across two joined balloons. Both detectors return one region spanning the join, and Telea smears the outline between the balloons into an X. The fill samples the ground colour at the region&apos;s text core and flood-fills pixels within 40 levels of it on every channel. The components that reach the core, with their holes filled, form the balloon. The letters are the holes, while the outline, including the one at the join, connects to the page outside and is left alone. The clipped region is then painted flat with the balloon&apos;s colour. Handing a clipped mask back to Telea failed, because once the mask stops at the outline, the ring Telea fills from is the outline. A guard sends a region back to Telea when more than 5% of its ink is letter-like shapes the fill would leave behind. This guard came from a failure: the first version, which checked for ink lost from the middle of the region, also switched off the fill on the joined balloons it was meant for.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Text drawn on art is labelled, not erased.&lt;/strong&gt; When a region gets no balloon fill and its surroundings fail the flat-field test below, its lettering is left in place. The translation goes in a small solid label beside it, at most 16 px and three lines. This removed the star-burst smears Telea produced on sound effects over artwork. One case remains open. A sound effect with a white outline over grey art looks like a balloon to the fill, so it gets painted as a white silhouette. We measured five hand-built separators, and none of them split outlined lettering from real balloons. The principled fix is a second output channel on TextSeg that predicts whether text sits inside a balloon, trained from the region kinds the labels already carry (14,528 bubble and 8,515 sound-effect regions in the automatic tier alone).&lt;/p&gt;
&lt;h3&gt;7.4 Rendering: contrast by construction&lt;/h3&gt;
&lt;p&gt;Detector boxes around vertical Japanese are tall and narrow: 11 of 13 regions on a test page had aspect ratios between 0.15 and 0.86. In three cases the box was narrower than a single English word of its own translation at a 10 px font, so shrinking alone could never fit the text. The layout may therefore widen a line up to 2.5 times the box width, centre it, and flag the region as overflowed. It never truncates, because the translation is the only copy the reader gets. Growing a box into surrounding white space was rejected after a test in which the growth escaped a balloon into the page gutter, from 32 px to 203 px wide. A white balloon interior and a white page look the same to a brightness test. With balloon fill, text is now set inside the balloon&apos;s own largest interior rectangles. A translation is split across the lobes of a joined balloon only when no part would come out smaller than the whole text would in the main lobe.&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;caption plate&lt;/strong&gt; goes behind text that has no balloon. The trigger is the share of the region, plus a 6 px margin, covered by its two dominant tones on the source page. Lettering on any flat field is two-tone by construction, and artwork is not. On 133 regions from 17 pages, a cut at 0.73 caught 5 of the 9 text-on-art regions with no false alarms among 124 bubbles. Looser cuts added false alarms faster than hits (0.89 gave 8 of 9 hits but 47 false alarms). Two other candidates failed. Ring uniformity separated perfectly on the first page and collapsed on held-out pages, flagging 24 of 57 regions with only about 5 correct, because a box that hugs its balloon puts the outline inside the ring. Measuring on the erased image instead of the source raised false alarms from 0 to 5, because Telea had already replaced the evidence with its own guess.&lt;/p&gt;
&lt;p&gt;The plate takes the panel&apos;s own tone, and the ink is black or white, whichever contrasts more. With relative luminance $L$, the contrast against white ink is $1.05/(L+0.05)$ and against black ink $(L+0.05)/0.05$. The two are equal at&lt;/p&gt;
&lt;p&gt;$$
L^{*} = \sqrt{0.0525} - 0.05 \approx 0.1791, \qquad \text{contrast} = \frac{\sqrt{0.0525}}{0.05} \approx 4.58 ,
$$&lt;/p&gt;
&lt;p&gt;so the better of the two choices never falls below the WCAG AA level of 4.5:1 [37] for any tone. Over all 136 regions the worst case is 8.57:1. Semi-transparency was the one gap in this guarantee. Blending with Telea&apos;s output underneath dropped the worst case to 3.28:1 for grounds with luma between 89 and 141, so plates in that band are drawn opaque.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;8. Discussion&lt;/h2&gt;
&lt;h3&gt;8.1 What generalises&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Judge detectors where the pipeline operates.&lt;/strong&gt; In this study the fixed-threshold comparison was misleading every time it disagreed with the matched one. v1&apos;s &quot;+0.31 sound-effect recall&quot; was mostly extra boxes. Seed 0 and seed 2 of recipe 1 looked identical on sound effects at 0.25 but were the worst and best at the gate. For a comic-translation pipeline that edits pixels, our experience suggests publishing the tolerated spurious rate and comparing models at that rate. We expect the same to hold for other pipelines that overwrite their detections, but this study does not test that.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A pooled reference is biased toward the detectors that proposed it.&lt;/strong&gt; Half of v4&apos;s apparent false positives were real text that the baseline had never proposed. The fix is hand review of every candidate&apos;s flagged boxes, with the reference re-frozen before each comparison. As models improve, this matters more, because a better model finds more text that the reference does not yet contain. The 2026 revision of Manga109 [8] reports the same kinds of missing region.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Report variance before recipes.&lt;/strong&gt; At our scale, the seed moved results as much as any recipe change did. Three seeds with a page bootstrap cost roughly three GPU-hours per recipe on a free T4, less than the Colab hours spent on single runs whose conclusions later reversed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Diagnose labels by source, not by score.&lt;/strong&gt; Which finders agreed on a region predicted real text far better than any finder&apos;s confidence, and one finder&apos;s confidence was inverted within its own group. Diagnostics grouped by source should come before any threshold tuning of an automatic labeller.&lt;/p&gt;
&lt;h3&gt;8.2 How this was built&lt;/h3&gt;
&lt;p&gt;This was a side project, built in spare time over about a month, not a formal academic study. We still tried to document the measurements carefully enough that the numbers can be checked. LLM coding agents (Claude Code) did a large share of the work: they implemented the extension and the training code, planned and ran experiments, analysed and audited results, and drafted this write-up. The author reviewed the reported measurements, did the hand annotation and review, made the technical and product decisions, and edited the final text. That review was mostly reading and spot-checking, not an independent re-run of every number, so errors of the kind the audits below caught may remain.&lt;/p&gt;
&lt;p&gt;Two habits mattered for the results. First, later sessions audited earlier ones against the raw training logs. That audit found the per-seed validation split, the coarse-grid artefact behind seed 0&apos;s apparent gate failure, the unsupervised-padding slivers, and the evidence against the box-size hypothesis (Section 6.3). This was separate from an earlier tooling bug found during the v4b run itself: the sweep script&apos;s &lt;code&gt;--out&lt;/code&gt; option overwrote the saved rows of models not named on the command line, and those rows were restored from version control before any comparison was made. Second, agent-reported numbers were treated as claims until reproduced. Trained models were re-scored locally from the files downloaded back from Hugging Face or Kaggle. Wherever a run had logged its own scores, the local numbers matched them.&lt;/p&gt;
&lt;h3&gt;8.3 Limitations&lt;/h3&gt;
&lt;p&gt;The reference has 57 pages, and the tight gate rests on two or three boxes, so we report bootstrap probabilities rather than claim significance. All pages come from one source, Danbooru, whose tag vocabulary shaped both sampling and slicing. The reviewing was done by one person without a second annotator, so labelling conventions (for example, excluding watermarks and printed cards from what needs translation) are one person&apos;s. TextSeg&apos;s ignore mask is untested by ablation, and its Japanese boxes remain about 1.5 times the reference area even with a tightened decoder. Browser detection differs slightly from the Python evaluation, because Chrome&apos;s canvas resize is not PIL&apos;s (per-region scores differ by up to 0.16). The translation comparison uses 23 lines and one annotator. On-device translation has not been re-verified since the pipeline moved to the offscreen document. Tiled webtoon strips are translated one tile at a time, not stitched.&lt;/p&gt;
&lt;h3&gt;8.4 Licensing and data ethics&lt;/h3&gt;
&lt;p&gt;Weights trained with Ultralytics&apos; code are covered, in that company&apos;s stated view, by AGPL-3.0, whatever licence tag a derivative carries. That was one reason for writing TextSeg. The released extension ships only TextSeg (Apache-2.0 backbone), manga-ocr and PP-OCR (Apache-2.0), OpenCV.js (Apache-2.0) and onnxruntime-web (MIT). Gemma is downloaded by the user&apos;s browser under the Gemma Terms of Use. The repositories store post IDs and boxes, never pages. Every page is re-fetched from Danbooru, whose per-image licence can be checked at fetch time. We avoided sources such as scanned doujinshi, whose rights status is unclear. The public repository was assembled from scratch, with no shared git history, after a full-history secret scan of the private one.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;9. Conclusion&lt;/h2&gt;
&lt;p&gt;We set out to put a comic translator into the browser and found that its hardest problem was measurement. A pipeline that erases and rewrites what it detects has a false-positive budget. For our baseline detector that budget was 0.9%, and at that budget most apparent improvements disappeared. Matched-ceiling sweeps, a reference corrected for pooling bias, and seed-and-page uncertainty turned a sequence of confusing reversals into a consistent account. Labels mattered and the starting checkpoint did not. Hard negatives were the first lever that moved the curve. Single runs could not rank recipes. With these methods in place, TextSeg, a 12.8 MB detector we wrote ourselves with an ignore-aware loss, beats the baseline at the tight ceiling on every seed, on one 57-page reference. It roughly triples sound-effect recall (0.41 against 0.15) and raises dark-background recall from 0.38 to 0.53.&lt;/p&gt;
&lt;p&gt;Three directions follow from what remains open. &lt;strong&gt;A balloon-versus-sound-effect channel&lt;/strong&gt; in TextSeg would let the pipeline fill real balloons and label outlined sound effects, removing the last major erase failure. &lt;strong&gt;A larger, independently reviewed reference&lt;/strong&gt;, extended by pooling every candidate&apos;s proposals, would narrow the bootstrap intervals enough to test the ignore mask by ablation. &lt;strong&gt;Calibrating the default threshold per language&lt;/strong&gt;, and moving detection and OCR to WebGPU, would make the matched-ceiling result the configuration users actually run. We think the protocol matters more than any single number in this paper. It was designed for, and tested on, one comic-translation pipeline; we expect it to be useful wherever a detector&apos;s errors end up drawn on the page, but showing that would take other datasets and pipelines.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;p&gt;[1] K. Aizawa, A. Fujimoto, A. Otsubo, T. Ogawa, Y. Matsui, K. Tsubota, H. Ikuta. &quot;Building a Manga Dataset &apos;Manga109&apos; with Annotations for Multimedia Applications.&quot; &lt;em&gt;IEEE MultiMedia&lt;/em&gt;, 27(2), 2020.&lt;/p&gt;
&lt;p&gt;[2] Y. Matsui, K. Ito, Y. Aramaki, A. Fujimoto, T. Ogawa, T. Yamasaki, K. Aizawa. &quot;Sketch-based Manga Retrieval using Manga109 Dataset.&quot; &lt;em&gt;Multimedia Tools and Applications&lt;/em&gt;, 76, 2017.&lt;/p&gt;
&lt;p&gt;[3] R. Hinami, S. Ishiwatari, K. Yasuda, Y. Matsui. &quot;Towards Fully Automated Manga Translation.&quot; &lt;em&gt;AAAI&lt;/em&gt;, 35(14):12998–13008, 2021. arXiv:2012.14271.&lt;/p&gt;
&lt;p&gt;[4] R. Sachdeva, A. Zisserman. &quot;The Manga Whisperer: Automatically Generating Transcriptions for Comics.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2024. arXiv:2401.10224.&lt;/p&gt;
&lt;p&gt;[5] J. Baek, Y. Matsui, K. Aizawa. &quot;COO: Comic Onomatopoeia Dataset for Recognizing Arbitrary or Truncated Texts.&quot; &lt;em&gt;ECCV&lt;/em&gt;, 2022. arXiv:2207.04675.&lt;/p&gt;
&lt;p&gt;[6] J. Del Gobbo, R. Matuk Herrera. &quot;Unconstrained Text Detection in Manga: A New Dataset and Baseline.&quot; &lt;em&gt;ECCV Workshops&lt;/em&gt;, 2020. arXiv:2009.04042.&lt;/p&gt;
&lt;p&gt;[7] E. Vivoli et al. &quot;One Missing Piece in Vision and Language: A Survey on Comics Understanding.&quot; arXiv:2409.09502, 2024.&lt;/p&gt;
&lt;p&gt;[8] &quot;Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding.&quot; &lt;em&gt;ICML Culture × AI Workshop&lt;/em&gt;, 2026. arXiv:2605.21182.&lt;/p&gt;
&lt;p&gt;[9] M. Liao, Z. Wan, C. Yao, K. Chen, X. Bai. &quot;Real-time Scene Text Detection with Differentiable Binarization.&quot; &lt;em&gt;AAAI&lt;/em&gt;, 2020.&lt;/p&gt;
&lt;p&gt;[10] mayocream. &lt;em&gt;koharu&lt;/em&gt; (software). https://github.com/mayocream/koharu&lt;/p&gt;
&lt;p&gt;[11] A. Howard et al. &quot;Searching for MobileNetV3.&quot; &lt;em&gt;ICCV&lt;/em&gt;, 2019.&lt;/p&gt;
&lt;p&gt;[12] T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, S. Belongie. &quot;Feature Pyramid Networks for Object Detection.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2017.&lt;/p&gt;
&lt;p&gt;[13] A. Shrivastava, A. Gupta, R. Girshick. &quot;Training Region-based Object Detectors with Online Hard Example Mining.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2016.&lt;/p&gt;
&lt;p&gt;[14] F. Milletari, N. Navab, S.-A. Ahmadi. &quot;V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation.&quot; &lt;em&gt;3DV&lt;/em&gt;, 2016.&lt;/p&gt;
&lt;p&gt;[15] P. Felzenszwalb, R. Girshick, D. McAllester, D. Ramanan. &quot;Object Detection with Discriminatively Trained Part-Based Models.&quot; &lt;em&gt;IEEE TPAMI&lt;/em&gt;, 32(9), 2010.&lt;/p&gt;
&lt;p&gt;[16] J. Redmon, S. Divvala, R. Girshick, A. Farhadi. &quot;You Only Look Once: Unified, Real-Time Object Detection.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2016.&lt;/p&gt;
&lt;p&gt;[17] A. Wang et al. &quot;YOLOv10: Real-Time End-to-End Object Detection.&quot; &lt;em&gt;NeurIPS&lt;/em&gt;, 2024.&lt;/p&gt;
&lt;p&gt;[18] Y. Zhao et al. &quot;DETRs Beat YOLOs on Real-time Object Detection.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2024.&lt;/p&gt;
&lt;p&gt;[19] W. Lv, Y. Zhao, Q. Chang, K. Huang, G. Wang, Y. Liu. &quot;RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer.&quot; arXiv:2407.17140, 2024.&lt;/p&gt;
&lt;p&gt;[20] Y. Du et al. &quot;PP-OCR: A Practical Ultra Lightweight OCR System.&quot; arXiv:2009.09941, 2020.&lt;/p&gt;
&lt;p&gt;[21] M. Li et al. &quot;TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models.&quot; &lt;em&gt;AAAI&lt;/em&gt;, 2023.&lt;/p&gt;
&lt;p&gt;[22] A. Dosovitskiy et al. &quot;An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.&quot; &lt;em&gt;ICLR&lt;/em&gt;, 2021.&lt;/p&gt;
&lt;p&gt;[23] A. Telea. &quot;An Image Inpainting Technique Based on the Fast Marching Method.&quot; &lt;em&gt;Journal of Graphics Tools&lt;/em&gt;, 9(1), 2004.&lt;/p&gt;
&lt;p&gt;[24] R. Suvorov et al. &quot;Resolution-robust Large Mask Inpainting with Fourier Convolutions.&quot; &lt;em&gt;WACV&lt;/em&gt;, 2022.&lt;/p&gt;
&lt;p&gt;[25] C. F. Ruan et al. &quot;WebLLM: A High-Performance In-Browser LLM Inference Engine.&quot; arXiv:2412.15803, 2024.&lt;/p&gt;
&lt;p&gt;[26] Gemma Team. &quot;Gemma 2: Improving Open Language Models at a Practical Size.&quot; arXiv:2408.00118, 2024.&lt;/p&gt;
&lt;p&gt;[27] C. Buckley, D. Dimmick, I. Soboroff, E. Voorhees. &quot;Bias and the Limits of Pooling for Large Collections.&quot; &lt;em&gt;Information Retrieval&lt;/em&gt;, 10:491–508, 2007.&lt;/p&gt;
&lt;p&gt;[28] J. Zobel. &quot;How Reliable Are the Results of Large-Scale Information Retrieval Experiments?&quot; &lt;em&gt;SIGIR&lt;/em&gt;, 1998.&lt;/p&gt;
&lt;p&gt;[29] X. Bouthillier et al. &quot;Accounting for Variance in Machine Learning Benchmarks.&quot; &lt;em&gt;MLSys&lt;/em&gt;, 2021.&lt;/p&gt;
&lt;p&gt;[30] B. Efron. &quot;Bootstrap Methods: Another Look at the Jackknife.&quot; &lt;em&gt;Annals of Statistics&lt;/em&gt;, 7(1), 1979.&lt;/p&gt;
&lt;p&gt;[31] B. T. Polyak, A. B. Juditsky. &quot;Acceleration of Stochastic Approximation by Averaging.&quot; &lt;em&gt;SIAM Journal on Control and Optimization&lt;/em&gt;, 30(4), 1992.&lt;/p&gt;
&lt;p&gt;[32] N. Otsu. &quot;A Threshold Selection Method from Gray-Level Histograms.&quot; &lt;em&gt;IEEE Transactions on Systems, Man, and Cybernetics&lt;/em&gt;, 9(1), 1979.&lt;/p&gt;
&lt;p&gt;[33] D.-H. Lee. &quot;Pseudo-Label: The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks.&quot; &lt;em&gt;ICML Workshop on Challenges in Representation Learning&lt;/em&gt;, 2013.&lt;/p&gt;
&lt;p&gt;[34] H. Song, M. Kim, D. Park, Y. Shin, J.-G. Lee. &quot;Learning from Noisy Labels with Deep Neural Networks: A Survey.&quot; &lt;em&gt;IEEE TNNLS&lt;/em&gt;, 2023.&lt;/p&gt;
&lt;p&gt;[35] A. Graves, S. Fernández, F. Gomez, J. Schmidhuber. &quot;Connectionist Temporal Classification.&quot; &lt;em&gt;ICML&lt;/em&gt;, 2006.&lt;/p&gt;
&lt;p&gt;[36] T. Fawcett. &quot;An Introduction to ROC Analysis.&quot; &lt;em&gt;Pattern Recognition Letters&lt;/em&gt;, 27(8), 2006.&lt;/p&gt;
&lt;p&gt;[37] W3C. &quot;Web Content Accessibility Guidelines (WCAG) 2.1.&quot; W3C Recommendation, 2018.&lt;/p&gt;
&lt;p&gt;[38] D. Smilkov et al. &quot;TensorFlow.js: Machine Learning for the Web and Beyond.&quot; &lt;em&gt;SysML&lt;/em&gt;, 2019.&lt;/p&gt;
&lt;p&gt;[39] X. Zhou et al. &quot;EAST: An Efficient and Accurate Scene Text Detector.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2017.&lt;/p&gt;
&lt;p&gt;[40] Y. Baek, B. Lee, D. Han, S. Yun, H. Lee. &quot;Character Region Awareness for Text Detection.&quot; &lt;em&gt;CVPR&lt;/em&gt;, 2019.&lt;/p&gt;
&lt;p&gt;[41] M. Everingham, L. Van Gool, C. Williams, J. Winn, A. Zisserman. &quot;The PASCAL Visual Object Classes (VOC) Challenge.&quot; &lt;em&gt;IJCV&lt;/em&gt;, 88(2), 2010.&lt;/p&gt;
&lt;p&gt;[42] Kaggle. &quot;Notebooks&quot; documentation (background &quot;Save &amp;#x26; Run All&quot; execution, GPU quotas). https://www.kaggle.com/docs/notebooks&lt;/p&gt;
&lt;p&gt;[43] D. Hasler, S. Süsstrunk. &quot;Measuring Colourfulness in Natural Images.&quot; &lt;em&gt;Proc. SPIE Human Vision and Electronic Imaging&lt;/em&gt;, 2003.&lt;/p&gt;
&lt;p&gt;[44] D. Picard. &quot;torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision.&quot; arXiv:2109.08203, 2021.&lt;/p&gt;
&lt;p&gt;[45] G. Jocher et al. &lt;em&gt;Ultralytics YOLO&lt;/em&gt; (software, AGPL-3.0). https://github.com/ultralytics/ultralytics&lt;/p&gt;
&lt;p&gt;[46] kha-white. &lt;em&gt;manga-ocr&lt;/em&gt; (software and model, Apache-2.0). https://github.com/kha-white/manga-ocr&lt;/p&gt;
&lt;p&gt;[47] ogkalu. &lt;em&gt;comic-text-and-bubble-detector&lt;/em&gt; (RT-DETR-v2 model, Apache-2.0). Hugging Face.&lt;/p&gt;
&lt;p&gt;[48] Kiuyha. &lt;em&gt;Manga-Bubble-YOLO&lt;/em&gt; (YOLO26n model). Hugging Face.&lt;/p&gt;
&lt;p&gt;[49] ogkalu. &lt;em&gt;comic-translate&lt;/em&gt; (software). https://github.com/ogkalu2/comic-translate&lt;/p&gt;
&lt;p&gt;[50] zyddnys. &lt;em&gt;manga-image-translator&lt;/em&gt; (software). https://github.com/zyddnys/manga-image-translator&lt;/p&gt;
&lt;p&gt;[51] Google. &quot;Manifest V3&quot; and &quot;chrome.offscreen&quot; developer documentation. https://developer.chrome.com/docs/extensions&lt;/p&gt;
&lt;p&gt;[52] Y. Aramaki, Y. Matsui, T. Yamasaki, K. Aizawa. &quot;Text Detection in Manga by Combining Connected-Component-Based and Region-Based Classifications.&quot; &lt;em&gt;ICIP&lt;/em&gt;, 2016.&lt;/p&gt;
&lt;p&gt;[53] T. Ogawa, A. Otsubo, R. Narita, Y. Matsui, T. Yamasaki, K. Aizawa. &quot;Object Detection for Comics using Manga109 Annotations.&quot; arXiv:1803.08670, 2018.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Appendix A. Artefacts and Reproducibility&lt;/h2&gt;
&lt;p&gt;:::details[Released artefacts]
| Artefact | Location |
|---|---|
| Extension source, MIT, TextSeg version | &lt;code&gt;vermilion10/comic-translator&lt;/code&gt; (GitHub), release v0.0.1 |
| TextSeg weights | Hugging Face &lt;code&gt;vermilion10/manga-textseg&lt;/code&gt; |
| YOLO fine-tunes v1, v2, v4, v5, v4b | Hugging Face &lt;code&gt;vermilion10/manga-bubble-yolo-finetune&lt;/code&gt; (tags) |
| YOLO v3 (COCO start) | Hugging Face &lt;code&gt;vermilion10/manga-text-yolo26&lt;/code&gt;, tag &lt;code&gt;v3&lt;/code&gt; |
| Detection reference (57 pages, 540 regions) | &lt;code&gt;ml/eval/detect_reference.json&lt;/code&gt; (post IDs and boxes) |
| OCR references | &lt;code&gt;ml/eval/{zh,ko,ja_page}_reference.json&lt;/code&gt; |
| Sweep and scoring | &lt;code&gt;ml/scripts/sweep_detect.py&lt;/code&gt;, &lt;code&gt;ml/eval/score_detect.py&lt;/code&gt; |
| Parity checks (Python ↔ TypeScript) | &lt;code&gt;extension/scripts/check-text-map.mjs&lt;/code&gt;, &lt;code&gt;check-balloon.mjs&lt;/code&gt; |
:::&lt;/p&gt;
&lt;p&gt;:::details[Training budgets]
| Run | Hardware | Epochs | Wall time |
|---|---|---|---|
| YOLO v1, v2, v4, v5, v4b | Colab T4 | 20 | 1.06–1.12 h each |
| YOLO v3 | Colab T4 | 40 | 2.19 h |
| TextSeg recipe 1, per seed | Kaggle T4 ×2 | 80 | 1.88–2.00 h |
| TextSeg recipe 2, three seeds | Kaggle T4 ×2 | 40 | 3.22 h total |
| TextSeg recipe 3, three seeds | Kaggle T4 ×2 | 40 | 3.04 h total |
:::&lt;/p&gt;
&lt;p&gt;:::details[Timeline (2026)]&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Aug 26–28.&lt;/strong&gt; Proposal, extension scaffold, then detect, OCR, erase, translate and render wired stage by stage.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Aug 28–31.&lt;/strong&gt; Page integration, host permissions, options, Chinese and Korean OCR, deduplication, floating control, manual boxes, caption plates, and the 57-page detection reference.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sep 2–3.&lt;/strong&gt; Training set, v1, the threshold sweep, the corroboration relabel, v2, the fragmentation diagnosis.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sep 10–12.&lt;/strong&gt; Cloud fallback, vertical Chinese OCR, release audit, v3, false-positive taxonomy, hard-negative mining.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sep 13–22.&lt;/strong&gt; v4, reference correction, second mining round, expansion (v5), seed replicate (v4b).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sep 23–26.&lt;/strong&gt; TextSeg design and three recipes on Kaggle, map-shaped erasing, balloon fill, Japanese chunking, offscreen pipeline, public release.
:::&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;[^1]: The reference grew from 525 to 540 regions as corrections were added, and every comparison was made on the reference as it stood at the time. Each detection table names its reference (Table 2: 525, Table 3: 535, Table 4: 540), and all rows in a table share it. Numbers quoted in the text say which reference they come from when it is not the final 540-region one.&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/><atom:updated>2026-10-05T00:00:00.000Z</atom:updated></item><item><title>Building My Own Blog Editor with Tauri</title><link>https://pure.vermilion10.dev/blog/building-my-own-blog-editor-with-tauri</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/building-my-own-blog-editor-with-tauri</guid><description>A Windows and Android app for writing this blog: three editor views, R2 image uploads and one-tap publishing to GitHub.</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Writing a post for this blog used to go like this on my phone: open Termux, &lt;code&gt;cd&lt;/code&gt; all the way into the repo, &lt;code&gt;git pull&lt;/code&gt;, switch to a text editor, write, switch back to Termux, then &lt;code&gt;git add&lt;/code&gt;, &lt;code&gt;git commit&lt;/code&gt;, &lt;code&gt;git push&lt;/code&gt;. Images were their own adventure: edit them in Image Toolbox, upload them with RSAF, then type the R2 key into the markdown by hand and hope I didn&apos;t make a typo.&lt;/p&gt;
&lt;p&gt;That&apos;s a lot of steps just to write a few paragraphs. So I thought, why not make an app that does only this one job? This post was written in that app.&lt;/p&gt;
&lt;p&gt;:::carousel{aspect=&quot;16/10&quot; fit=&quot;cover&quot;}
&lt;img src=&quot;posts/2026/blog-editotlr/code.webp&quot; alt=&quot;Code view&quot;&gt;
&lt;img src=&quot;posts/2026/blog-editotlr/preview.webp&quot; alt=&quot;Preview&quot;&gt;
&lt;img src=&quot;posts/2026/blog-editotlr/visualview.webp&quot; alt=&quot;Visual view&quot;&gt;
:::&lt;/p&gt;
&lt;h2&gt;What it does&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Three views&lt;/strong&gt;: a markdown code editor, a visual (WYSIWYG) editor with a formatting bar, and a live preview that looks exactly like the real site.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Straight to the repo&lt;/strong&gt;: it lists the posts from GitHub, keeps drafts on the device, and publishing is one commit through the GitHub API. No git and no terminal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Frontmatter form&lt;/strong&gt;: title, dates, tags (with suggestions from other posts), language and password protection, so I don&apos;t have to hand-edit YAML.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Syntax cheat sheet&lt;/strong&gt; in a sidebar. Tap an item to insert it at the cursor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Image uploads to R2&lt;/strong&gt; with conversion settings in the style of Image Toolbox: format, lossy or lossless, quality, or a target file size.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A change gutter&lt;/strong&gt; like VS Code&apos;s, marking the lines I&apos;ve added or changed since the last publish.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Runs on Windows and Android&lt;/strong&gt;, with a phone layout (bottom navigation bar) and a tablet layout (rail plus side pane).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The stack&lt;/h2&gt;
&lt;p&gt;| Part | Choice |
|---|---|
| App shell | Tauri 2 (Rust, with WebView2 on Windows and the system WebView on Android) |
| UI | Vue 3, Vite, Tailwind 4, Material Design 3 |
| Code editor | CodeMirror 6 |
| Visual editor | ProseMirror with prosemirror-markdown |
| Preview | the blog&apos;s own markdown-it renderer (Shiki, Mermaid, KaTeX) |
| Image encoding | jSquash (WebP, AVIF, JPEG, PNG, OxiPNG) in a Web Worker |
| Storage | Cloudflare R2, signed on the Rust side with rusty-s3 |
| Secrets | Windows Credential Manager, Android Keystore |&lt;/p&gt;
&lt;p&gt;I picked Tauri over Electron mostly for size. The Windows installer is about 10 MB and the Android release APK is 23 MB. The same Vue frontend runs on both, and the Rust side handles whatever the web view shouldn&apos;t: secrets, signing requests, and a few Android-only things.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;flowchart LR
    UI[&quot;Vue frontend&amp;#x3C;br/&gt;(CodeMirror, ProseMirror, preview)&quot;] --&gt;|posts, commits| GH[GitHub API]
    UI --&gt;|invoke| RS[Rust core]
    RS --&gt;|presigned requests| R2[Cloudflare R2]
    RS --&gt; KS[&quot;Credential Manager /&amp;#x3C;br/&gt;Android Keystore&quot;]
    UI --&gt;|encode| W[Image worker]
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;The preview is the real renderer&lt;/h2&gt;
&lt;p&gt;A preview that&apos;s only &lt;em&gt;close&lt;/em&gt; to the site is useless, because the thing I want to catch is exactly where they differ. So the app doesn&apos;t have its own renderer. A small script copies the blog&apos;s &lt;code&gt;MarkdownRenderer.vue&lt;/code&gt; into the app before every dev run and build, so the preview runs the same markdown-it pipeline, with the same plugins and the same CSS, as the live site.&lt;/p&gt;
&lt;p&gt;There&apos;s a catch: the blog repo is private, and the editor repo is public. So the copied renderer is git-ignored and never lands in the editor repo; it&apos;s pulled from my local checkout at build time.&lt;/p&gt;
&lt;h2&gt;The hard part: a visual editor that doesn&apos;t wreck my markdown&lt;/h2&gt;
&lt;p&gt;This was the part that took the longest. The usual way a WYSIWYG markdown editor works is: parse the markdown into a document, let you edit it, then serialize the whole thing back to markdown. The problem is that the serializer writes markdown &lt;em&gt;its&lt;/em&gt; way. &lt;code&gt;*italic*&lt;/code&gt; becomes &lt;code&gt;_italic_&lt;/code&gt;, list markers change, blank lines move, and my custom &lt;code&gt;:::&lt;/code&gt; blocks come back mangled. Open a post, fix one typo, and the commit diff touches every line.&lt;/p&gt;
&lt;p&gt;The fix was to remember where every block came from. When a post is loaded, each top-level block in the ProseMirror document is mapped back to the exact slice of the source it was parsed from. When saving, a block that hasn&apos;t been touched is written back byte for byte from the original. Only the blocks I actually edited go through the serializer.&lt;/p&gt;
&lt;p&gt;To make sure it holds, I ran all my existing posts through an open-and-save round trip with no edits. All 13 came back byte-identical, across 1147 blocks.&lt;/p&gt;
&lt;p&gt;The blog&apos;s custom blocks (Mermaid diagrams, GitHub cards, carousels, embeds) show up in the visual editor as cards that render the real thing, with an edit button that opens the raw source.&lt;/p&gt;
&lt;h3&gt;Making the diagrams fast&lt;/h3&gt;
&lt;p&gt;The first version froze. Opening the visual view on the Mermaid examples post blocked the UI for about 15 seconds, because every diagram was rendered up front, and the preview re-rendered every diagram on every keystroke. Two changes fixed it:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;Mermaid cache&lt;/strong&gt; keyed by the diagram source, so a diagram that didn&apos;t change is never rendered again, and&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;lazy cards&lt;/strong&gt; that render only when they scroll into view, with a skeleton in their place until then.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Opening the visual view went from 15.6 s blocked to 0.6 s, and a preview re-render from 4.9 s to 0.14 s.&lt;/p&gt;
&lt;h2&gt;Images: R2 without the storage manager&lt;/h2&gt;
&lt;p&gt;Uploading is drag and drop (or a file picker on Android). The image dialog lets me pick the format, lossy or lossless, quality, max width, or a target size. The encoding runs in a Web Worker with jSquash, so the UI doesn&apos;t stall. For a target size it does a binary search over the quality: encode, check the size, go up or down. When I set 100 KB, the file in the bucket came out at 97 KB.&lt;/p&gt;
&lt;p&gt;The R2 secret never touches the web view. The frontend only asks the Rust side to &quot;put these bytes at this key&quot;, and Rust signs the request with rusty-s3 and sends it with reqwest. The key follows the same layout my old posts already use, &lt;code&gt;posts/&amp;#x3C;year&gt;/&amp;#x3C;slug&gt;/&lt;/code&gt;, and uploads never overwrite. If &lt;code&gt;cover.webp&lt;/code&gt; already exists, the new one becomes &lt;code&gt;cover.v2.webp&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;:::note
On Android, Tauri can&apos;t hand raw bytes from the web view to Rust, so there the image goes over as base64 inside JSON. Desktop keeps the raw path. I only found this when the first upload on the tablet failed with &quot;expected file as raw body&quot;.
:::&lt;/p&gt;
&lt;h2&gt;The change gutter&lt;/h2&gt;
&lt;p&gt;I wanted the little colored bars you get in VS Code: green for added lines, blue for changed, a red marker where lines were deleted. CodeMirror has a merge extension, but its chunks are character-based and too coarse, so one small edit could paint half a paragraph. I replaced it with a plain line-level Myers diff against the published version, and fuzz-tested it with 3000 random edits.&lt;/p&gt;
&lt;h2&gt;Keeping the token safe&lt;/h2&gt;
&lt;p&gt;Signing in takes a fine-grained GitHub token limited to the blog repo with Contents set to read and write. Nothing more.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;On Windows&lt;/strong&gt; the token and the R2 keys go into the Windows Credential Manager through the &lt;code&gt;keyring&lt;/code&gt; crate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;On Android&lt;/strong&gt; there&apos;s a small Kotlin plugin that encrypts each value with AES-256-GCM, under a key that lives in the Android Keystore. Only the ciphertext is stored in the app&apos;s private storage.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The R2 secret is also marked Rust-only: the frontend can save it, but can never read it back.&lt;/p&gt;
&lt;h2&gt;Android was its own adventure&lt;/h2&gt;
&lt;p&gt;The desktop app worked early. Android took a few more rounds:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Edge to edge.&lt;/strong&gt; The first build drew under the status bar and the gesture bar. The WebView does report their size as safe-area insets, so every screen now pads itself with &lt;code&gt;env(safe-area-inset-*)&lt;/code&gt;, and the status bar icons switch between light and dark with the theme.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The keyboard.&lt;/strong&gt; Newer Android no longer resizes edge-to-edge apps when the keyboard opens, so it covered the editor. The app now sizes itself to &lt;code&gt;visualViewport&lt;/code&gt;, hides the bottom bar while typing, and scrolls the cursor back into view.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The back gesture&lt;/strong&gt; closes the open dialog first, then goes back to the post list, and only then leaves the app.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A red herring.&lt;/strong&gt; At one point R2 and even GitHub looked unreachable. It was Android&apos;s background network restriction, which kicked in because the app wasn&apos;t in the foreground while I was testing it over USB.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The UI&lt;/h2&gt;
&lt;p&gt;The UI follows Material Design 3. The whole palette is generated from one seed color with &lt;code&gt;material-color-utilities&lt;/code&gt;: orange by default, changeable in the settings, or taken from the wallpaper on Android 12 and up. It&apos;s dark by default with a light theme toggle. The layout changes with the width: bottom navigation on phones, a rail with a single pane on small tablets, and a rail plus a side pane (posts, details or cheat sheet) next to the editor on big screens like the Legion Y700.&lt;/p&gt;
&lt;h2&gt;Where it&apos;s at&lt;/h2&gt;
&lt;p&gt;It works, I&apos;m using it, and this post went through it from the first line to the publish button. There are ideas left, like slash commands in the visual editor and scroll sync for the split view, but for now I&apos;m parking it here.&lt;/p&gt;
&lt;p&gt;The code is on GitHub (the editor only; the blog itself stays private):&lt;/p&gt;
&lt;p&gt;::github{repo=&quot;vermilion10/vermilion10-blog-editor&quot;}&lt;/p&gt;</content:encoded><h:img src="https://cdn.vermilion10.dev/posts/2026/building-my-own-blog-editor-with-tauri/cover.webp"/><enclosure url="https://cdn.vermilion10.dev/posts/2026/building-my-own-blog-editor-with-tauri/cover.webp"/><atom:updated>2026-09-29T00:00:00.000Z</atom:updated></item><item><title>Reverse Engineering the Shiny Colors Asset Pipeline</title><link>https://pure.vermilion10.dev/blog/reverse-engineering-shiny-colors-asset-pipeline</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/reverse-engineering-shiny-colors-asset-pipeline</guid><description>A technical record of the investigation behind shinymas-dl, from client reconnaissance and asset-map reconstruction to authenticated card-hash discovery and the final C# implementation.</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h1&gt;Reverse Engineering the Shiny Colors Asset Pipeline&lt;/h1&gt;
&lt;p&gt;This document records the reverse-engineering work carried out while developing &lt;code&gt;shinymas-dl&lt;/code&gt;, a .NET downloader and extractor for &lt;em&gt;THE IDOLM@STER SHINY COLORS&lt;/em&gt;. It is an internal research record rather than a usage guide. Values that identify or authenticate an account are intentionally omitted.&lt;/p&gt;
&lt;p&gt;The investigation began with a simple question: given that the game runs entirely in the browser, how much of its asset pipeline can be reconstructed from the client itself? The answer was considerably more than expected. The browser client exposed a complete asset catalogue, the naming and hashing rules for ordinary CDN objects, the local format used for text resources, and, after authenticated API work, the metadata needed to resolve most of the card assets that initially appeared to be unavailable.&lt;/p&gt;
&lt;p&gt;The final result was a C# implementation with separate catalogue, download, naming, extraction, and API layers. The full album crawl recorded 1,440 card hashes for 1,461 card IDs present in asset-map v442, with no unowned card among the covered characters left without a hash.&lt;/p&gt;
&lt;p&gt;::github{repo=&quot;vermilion10/shinymas-dl&quot;}&lt;/p&gt;
&lt;h2&gt;1. Scope and initial reconnaissance&lt;/h2&gt;
&lt;p&gt;The starting point was an existing family of game asset tools: &lt;code&gt;kc-dl&lt;/code&gt;, &lt;code&gt;exilium-dl&lt;/code&gt;, and &lt;code&gt;pgrasset&lt;/code&gt;. The goal for &lt;code&gt;shinymas-dl&lt;/code&gt; was not to reproduce the game client as a whole, but to establish a reliable archive pipeline with three properties:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;the upstream catalogue should be recoverable without a game installation;&lt;/li&gt;
&lt;li&gt;the raw mirror should be resumable and independent from the final extraction layout; and&lt;/li&gt;
&lt;li&gt;naming and organization should come from data shipped by the game wherever possible.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The first pass therefore avoided writing application code. The site, client bundle, CDN behaviour, catalogue format, and resource transformations were investigated first.&lt;/p&gt;
&lt;p&gt;The game entry page was reachable without a geographic restriction. The HTML referenced a small &lt;code&gt;env.js&lt;/code&gt; followed by the main application bundle. The client identified the API root, asset root, game identifier, and a flag indicating that its resource-processing layer was enabled. The main bundle was approximately 3.7 MB, and later inspection identified 1,109 lazy-loaded chunks.&lt;/p&gt;
&lt;p&gt;The useful distinction at this stage was between information that could be obtained from the public client and information that depended on an authenticated game session. The former became the basis of the downloader. The latter was investigated separately and only integrated once its behaviour had been demonstrated with a test account.&lt;/p&gt;
&lt;h2&gt;2. Reconstructing the asset catalogue&lt;/h2&gt;
&lt;p&gt;The most important early discovery was that the client did not hard-code the complete list of resources. Instead, it loaded an asset-map manifest and a set of versioned chunks.&lt;/p&gt;
&lt;p&gt;The live manifest used 129 chunks. On the build investigated in this session, the result was:&lt;/p&gt;
&lt;p&gt;| Category | Files |
|---|---:|
| &lt;code&gt;sounds/&lt;/code&gt; | 273,121 |
| &lt;code&gt;images/&lt;/code&gt; | 76,921 |
| &lt;code&gt;spine/&lt;/code&gt; | 17,163 |
| &lt;code&gt;json/&lt;/code&gt; | 10,804 |
| &lt;code&gt;movies/&lt;/code&gt; | 3,352 |
| &lt;code&gt;ae/&lt;/code&gt; | 2,524 |
| &lt;code&gt;particles/&lt;/code&gt; | 2,213 |
| &lt;code&gt;fonts/&lt;/code&gt; | 4 |
| &lt;strong&gt;Total&lt;/strong&gt; | &lt;strong&gt;386,102&lt;/strong&gt; |&lt;/p&gt;
&lt;p&gt;The manifest reported 35.19 GiB in total.&lt;/p&gt;
&lt;p&gt;Each chunk maps a logical asset path to a version value. Those paths are the names used by the application before the CDN-specific filename transformation is applied. That distinction mattered because the visible game resource path and the actual URL requested from the CDN are not the same thing.&lt;/p&gt;
&lt;p&gt;The refresh implementation keeps the manifest and chunks in &lt;code&gt;data/&lt;/code&gt;. A subsequent refresh compares chunk versions and downloads only chunks whose version changed. The manifest is written last, so an interrupted update does not leave a new manifest pointing at a partially updated chunk set.&lt;/p&gt;
&lt;p&gt;The first successful run reproduced all 386,102 entries locally. A second refresh against the unchanged manifest performed no unnecessary chunk downloads.&lt;/p&gt;
&lt;h2&gt;3. CDN naming and resource transformations&lt;/h2&gt;
&lt;h3&gt;3.1 Filename hashing&lt;/h3&gt;
&lt;p&gt;The browser does not request a resource by its logical path. The client-side hashing function takes the filename stem, its first and last characters, and the full path under &lt;code&gt;/assets/&lt;/code&gt;, then hashes the result with SHA-256.&lt;/p&gt;
&lt;p&gt;Conceptually, for an asset path &lt;code&gt;p&lt;/code&gt; with filename stem &lt;code&gt;s&lt;/code&gt;, the input is:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;first(s) + last(s) + &quot;/assets/&quot; + p
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The digest is represented as lowercase hexadecimal. Media files keep their &lt;code&gt;.mp3&lt;/code&gt;, &lt;code&gt;.mp4&lt;/code&gt;, or &lt;code&gt;.m4a&lt;/code&gt; extension on the CDN; the other resource classes use the bare hexadecimal name.&lt;/p&gt;
&lt;p&gt;This was not inferred from a single sample. The same rule reproduced hashes for images, Spine files, JSON, animation data, audio and video. A representative ordinary path is:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;images/bg/001.jpg
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;which becomes a 64-character hexadecimal object name derived from the rule above.&lt;/p&gt;
&lt;p&gt;The asset map therefore contains enough information to construct ordinary CDN URLs without consulting any external database.&lt;/p&gt;
&lt;h3&gt;3.2 Text resources&lt;/h3&gt;
&lt;p&gt;The next obstacle was that not all resources were served as their final plaintext representation. JSON and Spine &lt;code&gt;.atlas&lt;/code&gt; resources were transformed before being written to the network.&lt;/p&gt;
&lt;p&gt;Static inspection of the client revealed &lt;code&gt;picosha2&lt;/code&gt;, zlib, an embedded string used as a key, and an &lt;code&gt;inflate&lt;/code&gt; path. Several hypotheses were tested against known resource files. The useful signal was not a printable-text score but the compressed data itself: once the XOR transformation was correct, the resulting stream had a valid gzip/deflate structure and decompressed to the same plaintext used by the game.&lt;/p&gt;
&lt;p&gt;The implemented transformation is a repeating XOR followed by raw-deflate decompression. The embedded ASCII key is 54 bytes long. After XOR, the resource contains the gzip header and a deflate body; the client-side format omits the normal gzip trailer, so the C# port inflates the raw deflate portion directly.&lt;/p&gt;
&lt;p&gt;The important implementation decision was to keep this transformation in the raw-to-plain extraction boundary. The raw mirror remains a faithful copy of the game resources, while &lt;code&gt;extract&lt;/code&gt; is responsible for producing readable JSON, scripts and other derived outputs.&lt;/p&gt;
&lt;p&gt;For a concrete validation, a downloaded &lt;code&gt;json/produce_events/...json&lt;/code&gt; resource was decrypted and parsed successfully, and a large Spine &lt;code&gt;.json&lt;/code&gt; resource expanded to readable JSON. Images, audio, video and fonts did not require this transformation.&lt;/p&gt;
&lt;h2&gt;4. The CDN&apos;s most misleading property&lt;/h2&gt;
&lt;p&gt;Once the ordinary path transformation worked, the next problem appeared during bulk testing: a missing resource did not necessarily produce an HTTP 404.&lt;/p&gt;
&lt;p&gt;Unknown paths were served as the game&apos;s single-page application shell with status 200 and an HTML content type. A downloader that considered only the HTTP status would therefore save copies of &lt;code&gt;index.html&lt;/code&gt; under resource filenames.&lt;/p&gt;
&lt;p&gt;The downloader treats an HTML response as the CDN&apos;s SPA fallback. Missing entries are recorded in the download index so that a normal resumed run does not repeatedly probe the same unavailable object. A &lt;code&gt;--retry-missing&lt;/code&gt; option removes those records when a later investigation shows that the object may have become available.&lt;/p&gt;
&lt;p&gt;The distinction is useful beyond this project. The server&apos;s HTTP status was technically successful while the requested resource was absent. Content validation was therefore part of the downloader&apos;s correctness, not merely an error-handling detail.&lt;/p&gt;
&lt;h2&gt;5. Building the first downloader&lt;/h2&gt;
&lt;p&gt;The first implementation mirrored the asset map directly into &lt;code&gt;output/raw/&lt;/code&gt;. The logical path was preserved, which made the raw tree independent from any future naming convention.&lt;/p&gt;
&lt;p&gt;The download index records the upstream version and local size for each completed resource. A file is skipped when the version is unchanged and its on-disk size still matches the recorded value. A version change forces a refetch.&lt;/p&gt;
&lt;p&gt;Index writes are performed atomically. The implementation serializes a sorted snapshot into a temporary file and then moves it into place. Progress is stored periodically rather than only at the end, because a full mirror is too large for an interrupted run to be treated as a single transaction.&lt;/p&gt;
&lt;p&gt;Downloads use bounded parallelism through &lt;code&gt;Parallel.ForEachAsync&lt;/code&gt;. Large files are written to a temporary &lt;code&gt;.part&lt;/code&gt; file and renamed only after the transfer completes. Network failures and transient request failures are retried with bounded exponential backoff.&lt;/p&gt;
&lt;p&gt;One practical test downloaded the complete story-script category: 10,804 files totaling about 21 MiB, using a concurrency of 16. The run took roughly four minutes. The limiting factor at that size was request latency rather than aggregate bandwidth.&lt;/p&gt;
&lt;p&gt;The resulting CLI was split into five operations:&lt;/p&gt;
&lt;p&gt;| Command | Function |
|---|---|
| &lt;code&gt;refresh&lt;/code&gt; | Fetch and update the asset map |
| &lt;code&gt;list&lt;/code&gt; | Search the cached map and produce statistics |
| &lt;code&gt;download&lt;/code&gt; | Mirror resources into the raw tree |
| &lt;code&gt;names&lt;/code&gt; | Build character names from story data |
| &lt;code&gt;extract&lt;/code&gt; | Decrypt and organize the mirror offline |&lt;/p&gt;
&lt;p&gt;This separation proved useful later when the resource-availability assumptions changed. The downloader did not need to be rewritten simply because the extraction rules grew more complicated.&lt;/p&gt;
&lt;h2&gt;6. Naming the data without an external database&lt;/h2&gt;
&lt;h3&gt;6.1 The initial problem&lt;/h3&gt;
&lt;p&gt;The catalogue identifies characters and cards primarily by numeric IDs. A raw mirror such as&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;images/content/characters/icon_circle/001.png
images/content/idols/card/1040010010.jpg
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;is machine-readable but not particularly useful as an archive.&lt;/p&gt;
&lt;p&gt;The project therefore set a constraint: use the game&apos;s own data to build names instead of depending on a community-maintained character table.&lt;/p&gt;
&lt;h3&gt;6.2 Story scripts as a name source&lt;/h3&gt;
&lt;p&gt;The story resources were particularly useful because their records contain information about the character visible in a scene, a romanized character label, and the Japanese speaker name.&lt;/p&gt;
&lt;p&gt;After the text resources were decrypted, the implementation scanned the mirrored scripts and built a name index from repeated observations. The resulting cache contains 33 characters: all 28 idols, Hazuki, and four collaboration characters.&lt;/p&gt;
&lt;p&gt;The naming process is intentionally statistical rather than based on a single line. A scene may identify one character while another character is speaking, so the visible character label and speaker name are collected as separate observations and combined by frequency.&lt;/p&gt;
&lt;p&gt;The complete script collection was sufficient to build the table locally in about 1.2 seconds. No external character metadata was needed.&lt;/p&gt;
&lt;h3&gt;6.3 Card ID structure&lt;/h3&gt;
&lt;p&gt;The card IDs themselves also contained useful structure. A representative produce card ID can be separated as:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;1 | 04 | 001 | 0010
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;where the fields represent card type, rarity, character ID, and sequence within that character&apos;s card set. Support cards use the corresponding support-card type.&lt;/p&gt;
&lt;p&gt;This meant that most card paths could be organized directly from their own IDs. For example, character &lt;code&gt;001&lt;/code&gt; can be resolved to the &lt;code&gt;mano&lt;/code&gt; name through &lt;code&gt;names.json&lt;/code&gt;, and a card beginning &lt;code&gt;104001...&lt;/code&gt; can be routed without a separate card database.&lt;/p&gt;
&lt;p&gt;The structure also explained several apparently unusual groups beginning with &lt;code&gt;193&lt;/code&gt;, &lt;code&gt;194&lt;/code&gt;, &lt;code&gt;293&lt;/code&gt;, and &lt;code&gt;294&lt;/code&gt;. At first those prefixes did not match the normal rarity constants found in the client. They were later shown to be present in the album data as well, so they were not malformed IDs.&lt;/p&gt;
&lt;h2&gt;7. Extraction layout&lt;/h2&gt;
&lt;p&gt;The raw mirror deliberately preserves upstream naming. &lt;code&gt;extract&lt;/code&gt; is the layer that turns it into an archive organized for human use.&lt;/p&gt;
&lt;p&gt;The principal routes are:&lt;/p&gt;
&lt;p&gt;| Source | Extracted form |
|---|---|
| Character image | &lt;code&gt;characters/&amp;#x3C;id&gt;_&amp;#x3C;name&gt;/...&lt;/code&gt; |
| Idol card | &lt;code&gt;idols/&amp;#x3C;id&gt;_&amp;#x3C;name&gt;/&amp;#x3C;cardId&gt;/...&lt;/code&gt; |
| Support card | &lt;code&gt;support_idols/&amp;#x3C;id&gt;_&amp;#x3C;name&gt;/&amp;#x3C;cardId&gt;/...&lt;/code&gt; |
| Story JSON | &lt;code&gt;scenarios/&amp;#x3C;type&gt;/&amp;#x3C;id&gt;/script.json&lt;/code&gt; |
| Story voice | &lt;code&gt;scenarios/&amp;#x3C;type&gt;/&amp;#x3C;id&gt;/voice/&amp;#x3C;file&gt;.m4a&lt;/code&gt; |
| Everything else | &lt;code&gt;other/&amp;#x3C;original game path&gt;&lt;/code&gt; |&lt;/p&gt;
&lt;p&gt;A readable &lt;code&gt;script.txt&lt;/code&gt; is generated alongside story JSON. The text representation includes the speaker, line, choices where applicable, and a matching voice path when one exists.&lt;/p&gt;
&lt;p&gt;The extractor also handles Spine directories and the game&apos;s other structured resource types. Text assets are decrypted during extraction, so the raw mirror remains suitable for later reprocessing if the archive layout changes.&lt;/p&gt;
&lt;h2&gt;8. The first coverage estimate was wrong&lt;/h2&gt;
&lt;p&gt;At this point the downloader appeared mostly complete, but a random reachability test suggested that only about ten percent of resources were missing. That estimate was misleading because the sample was dominated by resource classes that remained publicly reachable, especially voice and Spine assets.&lt;/p&gt;
&lt;p&gt;A full audit of card-related paths gave a different result:&lt;/p&gt;
&lt;p&gt;| Resource group | Reachable without additional metadata |
|---|---:|
| Produce card images | 68 / 546 |
| Support card images | 76 / 915 |
| SSR produce card images | 19 / 348 |
| Produce card icons | same per-card reachability as card art |
| Whisper voice files | 0 / 42 |&lt;/p&gt;
&lt;p&gt;The exact equivalence between idol card art and idol icons was significant. The same cards failed in the same pattern, including the same rarity distribution. That strongly suggested that the problem was attached to a card-level attribute rather than to individual files.&lt;/p&gt;
&lt;p&gt;The investigation then shifted from the CDN to the client code that constructs resource paths.&lt;/p&gt;
&lt;h2&gt;9. Discovering the protected asset convention&lt;/h2&gt;
&lt;p&gt;Static analysis of the main bundle and its lazy-loaded chunks identified a second URL form used by card resources:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;&amp;#x3C;hash&gt;_&amp;#x3C;id&gt;.&amp;#x3C;extension&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The hash was not derived from the individual file path. It belonged to the card record itself.&lt;/p&gt;
&lt;p&gt;The client uses the same card hash for several related resources, including:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;card icons;&lt;/li&gt;
&lt;li&gt;card art;&lt;/li&gt;
&lt;li&gt;card thumbnails;&lt;/li&gt;
&lt;li&gt;some card movies;&lt;/li&gt;
&lt;li&gt;awake and memorial-related artwork;&lt;/li&gt;
&lt;li&gt;certain fes-related images.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The code also exposed several hash-bearing fields with slightly different roles:&lt;/p&gt;
&lt;p&gt;| Field | Observed use |
|---|---|
| &lt;code&gt;hash&lt;/code&gt; | Main card resource paths |
| &lt;code&gt;idolHash&lt;/code&gt; | Awake and memorial-related paths |
| &lt;code&gt;supportIdolHash&lt;/code&gt; | Support-card resources |
| &lt;code&gt;centerIdolHash&lt;/code&gt; | Fes icon and state-specific resources |
| &lt;code&gt;voiceHash&lt;/code&gt; | Individual whisper voice files |&lt;/p&gt;
&lt;p&gt;Other game systems have their own independent hashes, such as gasha resources and some event-specific content. Those were outside the card-asset scope of this investigation.&lt;/p&gt;
&lt;p&gt;The asset map was therefore not actually missing the card objects. It was missing one piece of information needed to construct their current CDN object names.&lt;/p&gt;
&lt;h2&gt;10. Where the card hashes came from&lt;/h2&gt;
&lt;p&gt;The character album turned out to be the key data source.&lt;/p&gt;
&lt;p&gt;Static analysis found a character album request of the form:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;POST characterAlbums/characters/{characterId}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The album entries contain card IDs and their hash values. The client also tracks ownership separately through fields such as &lt;code&gt;hasIdol&lt;/code&gt; and &lt;code&gt;hasSupportIdol&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;A particularly important observation came from the album UI code. The album constructs an icon for every entry, not only the cards owned by the account. Unowned cards are rendered as grey and are made non-interactive, but they are still part of the list. The card-art screen is only entered for owned cards.&lt;/p&gt;
&lt;p&gt;That resolved a confusion from the original network captures. The apparent &quot;availability after a condition&quot; was not evidence that the server omitted every piece of metadata until the condition was met. The more precise description was that the UI withheld the art request behind an ownership condition while the album data already contained the card record and its hash.&lt;/p&gt;
&lt;p&gt;The list of characters displayed by the album comes from &lt;code&gt;GET album/top&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;11. Understanding the authenticated API transport&lt;/h2&gt;
&lt;p&gt;The game&apos;s API uses a different transport from ordinary CDN resources. The browser serializes the logical request as a raw HTTP request, encrypts that representation, and sends it as a single &lt;code&gt;application/octet-stream&lt;/code&gt; POST.&lt;/p&gt;
&lt;p&gt;The relevant client module is named &lt;code&gt;request_hash&lt;/code&gt;. Static analysis showed an asm.js build and a WebAssembly variant. The surrounding code identifies the following behaviour:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Construct the raw request text, including method, path, headers, and body.&lt;/li&gt;
&lt;li&gt;Pass the text through &lt;code&gt;encodeRequest&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Send the resulting bytes to the game API as &lt;code&gt;application/octet-stream&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Receive an encrypted response body as an array buffer.&lt;/li&gt;
&lt;li&gt;Decode the response with the &lt;code&gt;x-sessionid&lt;/code&gt; returned by the server.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The API client uses a rotating session identifier. Requests are queued and sent serially because the response to one request supplies the session identifier used by the next request.&lt;/p&gt;
&lt;p&gt;The browser build investigated in this session used &lt;code&gt;x-version: 410&lt;/code&gt; and an &lt;code&gt;x-em-version&lt;/code&gt; value embedded in the client bundle. The API client also recognised retryable statuses &lt;code&gt;1012&lt;/code&gt; and &lt;code&gt;1090&lt;/code&gt;, and it stops when the response contains &lt;code&gt;x-is-banned&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The request codec was first exercised offline. A request encoded with the game&apos;s own module could be decoded back to its original text. This demonstrated that the format itself was understood before any authenticated API experiment was attempted.&lt;/p&gt;
&lt;h3&gt;11.1 A misleading transport failure&lt;/h3&gt;
&lt;p&gt;Early attempts from Node produced HTTP 403 responses. The initial assumption was that the environment might be geographically restricted or that the API required a special browser fingerprint.&lt;/p&gt;
&lt;p&gt;A controlled test disproved both explanations. The same encoded no-token login was accepted by the game server when sent with &lt;code&gt;curl&lt;/code&gt;, producing HTTP 401 and an encrypted error response. .NET&apos;s &lt;code&gt;HttpClient&lt;/code&gt;, both with its default User-Agent and with a browser-style User-Agent, also reached the server and produced the same 401.&lt;/p&gt;
&lt;p&gt;The resulting decoded error was:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{&quot;title&quot;:&quot;Unauthorized&quot;,&quot;status&quot;:1010,&quot;instance&quot;:&quot;/login&quot;}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The 403 therefore belonged to the earlier transport path rather than to a game-side geographic block. This mattered because it prevented the API implementation from being built around a false region-lock theory.&lt;/p&gt;
&lt;h2&gt;12. The request codec implementation&lt;/h2&gt;
&lt;p&gt;One could fully reimplement the game&apos;s request codec in native C#, but there was no practical reason to do that for the downloader. The game already supplies the codec module on its public client side, and the client changes whenever the game changes.&lt;/p&gt;
&lt;p&gt;The implementation instead resolves the live client at runtime. It reads the current &lt;code&gt;index.html&lt;/code&gt;, &lt;code&gt;env.js&lt;/code&gt;, and application bundle, extracts the API root, game ID, client version headers, and the current &lt;code&gt;request_hash&lt;/code&gt; chunk name, then caches those client files under &lt;code&gt;data/client/&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The codec is executed inside .NET using Jint. The implementation loads the game&apos;s own JavaScript module and calls its exported &lt;code&gt;encodeRequest&lt;/code&gt; and &lt;code&gt;decodeResponse&lt;/code&gt; functions. No copy of the game&apos;s codec source is bundled into the repository.&lt;/p&gt;
&lt;p&gt;This decision has two advantages. First, the tool remains aligned with the current client bundle. Second, it avoids maintaining a native implementation of a format whose details are internal to the game.&lt;/p&gt;
&lt;p&gt;The separate reverse-engineering experiments did, however, reveal a useful property of the request format. An encoded request has an eight-byte header, and the output changes with the current Unix second. Holding the input constant and delaying the call changed the encoded output. Static inspection tied the timestamp source to &lt;code&gt;Date.now()&lt;/code&gt; and the native implementation to an SFMT-based generator. That observation was enough to explain why repeated identical requests do not necessarily produce identical ciphertext.&lt;/p&gt;
&lt;h2&gt;13. Establishing the login chain&lt;/h2&gt;
&lt;p&gt;The remaining problem was obtaining a valid game session for the authenticated album calls.&lt;/p&gt;
&lt;p&gt;The browser&apos;s Enza integration first supplies a short-lived game token, which is then used by the game login endpoint. A successful login returns the initial game &lt;code&gt;X-Sessionid&lt;/code&gt;; subsequent API calls rotate this value.&lt;/p&gt;
&lt;p&gt;The test sequence that finally succeeded was:&lt;/p&gt;
&lt;p&gt;| Step | Request | Result |
|---|---|---|
| 1 | Enza platform token exchange using the browser session | Short-lived game token |
| 2 | Game &lt;code&gt;POST /login&lt;/code&gt; with the game token | Game login succeeded; initial session established |
| 3 | &lt;code&gt;GET gameTop&lt;/code&gt; | Session headers returned |
| 4 | &lt;code&gt;GET album/top&lt;/code&gt; | 28 album characters |
| 5 | &lt;code&gt;POST characterAlbums/characters/1&lt;/code&gt; | Complete character album |&lt;/p&gt;
&lt;p&gt;No account credential is stored in the project. The session value is supplied to the process environment for the duration of a run and is not written to the data directory, logs, source tree, generated JSON, or git history.&lt;/p&gt;
&lt;p&gt;One false start came from using a different Enza token type. The platform treats those credentials differently. The game accepted the correct browser session-derived game token but treated the other token as a guest context. This was eventually resolved by inspecting the actual browser session flow rather than assuming that every Enza token represented the same credential layer.&lt;/p&gt;
&lt;h2&gt;14. The first authenticated album test&lt;/h2&gt;
&lt;p&gt;Character &lt;code&gt;001&lt;/code&gt; was used as the controlled test case before the full crawl.&lt;/p&gt;
&lt;p&gt;The album contained 64 cards: 22 produce cards and 42 support cards. The test account owned only a small subset of them. Nevertheless, every one of the 64 album entries carried a non-empty 32-character hash, including the cards the account did not own.&lt;/p&gt;
&lt;p&gt;The album also contained 23 costume entries with their own hash fields.&lt;/p&gt;
&lt;p&gt;The important end-to-end test used the unowned card &lt;code&gt;1940010010&lt;/code&gt;. The ordinary path continued to return the SPA fallback, while the hashed CDN form returned a real JPEG with a size of 76,235 bytes. This established the entire chain:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;asset map card ID
        |
        v
character album
        |
        v
card hash
        |
        v
&amp;#x3C;hash&gt;_&amp;#x3C;cardId&gt;.jpg
        |
        v
real CDN image
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This was the point at which the original reachability problem stopped being a catalogue problem. The downloader had already known that the card existed. It now had a reproducible mapping from the logical card ID to the current protected CDN object name.&lt;/p&gt;
&lt;h2&gt;15. Whisper voices and client-side release state&lt;/h2&gt;
&lt;p&gt;Whisper voices required a separate check because they use &lt;code&gt;voiceHash&lt;/code&gt; rather than the main card hash.&lt;/p&gt;
&lt;p&gt;The support-card data contains a list of whisper voice records together with their release conditions. The client constructs the voice path from &lt;code&gt;idolId&lt;/code&gt;, voice ID, and &lt;code&gt;voiceHash&lt;/code&gt;, then evaluates whether the voice is released from the card&apos;s evolution stage.&lt;/p&gt;
&lt;p&gt;Character &lt;code&gt;014&lt;/code&gt; provided the decisive test case. The test account owned none of the support cards carrying the relevant whisper set. Six voice entries were nevertheless present, all marked as locked by the client and all carrying &lt;code&gt;voiceHash&lt;/code&gt; values.&lt;/p&gt;
&lt;p&gt;All six downloaded through their hashed paths. Together they accounted for 64.6 MiB.&lt;/p&gt;
&lt;p&gt;The observed model is therefore slightly different from the initial assumption that a locked voice simply did not exist to the client. The client receives both the resource metadata and the release condition, then decides locally which tracks should be exposed in the current UI state.&lt;/p&gt;
&lt;p&gt;Six whisper files in the asset map remained unresolved. They correspond to the &lt;code&gt;...000.m4a&lt;/code&gt; track in each set and are not referenced by an album entry in the data examined during this session.&lt;/p&gt;
&lt;h2&gt;16. Full album crawl&lt;/h2&gt;
&lt;p&gt;After the controlled tests passed, the same process was applied to all 28 characters returned by &lt;code&gt;album/top&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The crawl completed in 59.4 seconds on the test environment.&lt;/p&gt;
&lt;p&gt;| Result | Count |
|---|---:|
| Album characters processed | 28 |
| Card hashes recorded | 1,440 |
| Card IDs in asset map v442 | 1,461 |
| Card IDs without a hash | 21 |
| Unowned cards without a hash among album characters | 0 |
| Whisper voices with &lt;code&gt;voiceHash&lt;/code&gt; | 36 |
| Whisper files in asset map | 42 |&lt;/p&gt;
&lt;p&gt;The 21 cards without a recorded hash were concentrated in characters not represented by the album list used by the crawler: the four collaboration characters and Hazuki. The six unresolved whisper files were the &lt;code&gt;...000&lt;/code&gt; member of each whisper set rather than normal album-referenced voice entries.&lt;/p&gt;
&lt;p&gt;The crawl was also made resumable. A completed character is recorded in the local album-hash store, so a second invocation skipped all 28 characters rather than repeating the API calls.&lt;/p&gt;
&lt;p&gt;This changed the downloader&apos;s semantics from &quot;try every known CDN path&quot; to &quot;build the best-known current CDN path from game metadata, then fall back to the ordinary path when no hash is known.&quot;&lt;/p&gt;
&lt;h2&gt;17. Integrating the album data into &lt;code&gt;shinymas-dl&lt;/code&gt;&lt;/h2&gt;
&lt;p&gt;The resulting feature was implemented as an &lt;code&gt;albums&lt;/code&gt; command.&lt;/p&gt;
&lt;p&gt;The command performs the following operations:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;resolve the current client bundle and API transport;&lt;/li&gt;
&lt;li&gt;obtain an authenticated game session from the supplied environment credential;&lt;/li&gt;
&lt;li&gt;read the album character list;&lt;/li&gt;
&lt;li&gt;fetch each character album serially;&lt;/li&gt;
&lt;li&gt;store card and costume IDs with their corresponding hashes; and&lt;/li&gt;
&lt;li&gt;make those hashes available to the download layer.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The persistent album file contains only resource IDs and hashes. It does not contain authentication state.&lt;/p&gt;
&lt;p&gt;The downloader&apos;s asset representation was extended with an optional &lt;code&gt;CdnPath&lt;/code&gt;. If a card hash is known, the hashed CDN form is attempted first. If no hash is available, the existing ordinary path logic remains unchanged. This preserves the original behaviour for resources outside the album-derived set.&lt;/p&gt;
&lt;p&gt;The API-side client also follows the browser&apos;s sequential session behaviour. It refreshes the current session ID after each response, applies the observed retry semantics, and treats the ban header as a hard stop rather than trying to continue.&lt;/p&gt;
&lt;h2&gt;18. Validation and security checks&lt;/h2&gt;
&lt;p&gt;Because the album feature is the only part of the project that handles an account session, the implementation was checked separately from the functional tests.&lt;/p&gt;
&lt;p&gt;The following checks were performed after the completed crawl:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;the session value did not appear in git history;&lt;/li&gt;
&lt;li&gt;no JWT-looking token strings appeared in git history;&lt;/li&gt;
&lt;li&gt;the &lt;code&gt;data/&lt;/code&gt; and &lt;code&gt;output/&lt;/code&gt; trees were not tracked by git;&lt;/li&gt;
&lt;li&gt;the card-hash JSON contained only IDs and hashes;&lt;/li&gt;
&lt;li&gt;the session value, game token, and per-request session IDs were not printed or persisted by the API client.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The browser session was treated as a credential and was not copied into the research document. The test account should be logged out of Enza after experimentation rather than leaving the session active indefinitely.&lt;/p&gt;
&lt;h2&gt;19. Bugs and corrections&lt;/h2&gt;
&lt;p&gt;The investigation also produced several failures unrelated to the asset protection itself. They were important because each one could have produced a plausible but incorrect archive.&lt;/p&gt;
&lt;h3&gt;19.1 The name-cache regression&lt;/h3&gt;
&lt;p&gt;The original &lt;code&gt;extract&lt;/code&gt; command rebuilt &lt;code&gt;names.json&lt;/code&gt; from whatever story scripts happened to be present in the current raw mirror. On a deliberately small test mirror, that reduced the source set to two scripts and replaced the previously complete 33-entry cache with a one-entry table.&lt;/p&gt;
&lt;p&gt;The failure was silent because the extraction still completed successfully.&lt;/p&gt;
&lt;p&gt;The fix was to merge a newly built name table into the cached table. Newly observed values replace stale entries, while names not represented in the current sample remain in the cache.&lt;/p&gt;
&lt;p&gt;This is a general issue with derived caches: a smaller input set is not evidence that previously known data should be deleted.&lt;/p&gt;
&lt;h3&gt;19.2 Awake images initially routed as generic output&lt;/h3&gt;
&lt;p&gt;Awake-step images were initially falling under &lt;code&gt;other/&lt;/code&gt; because their resource paths did not match the first set of extraction rules. The underlying card association was later shown by the client to use the card hash, but the extraction routing itself remained separate from the CDN lookup. This is an organizational issue rather than a download failure.&lt;/p&gt;
&lt;h3&gt;19.3 The progress renderer&lt;/h3&gt;
&lt;p&gt;The terminal progress renderer used carriage returns to update a single line. When stdout was redirected, those updates became permanent log lines. The fix was to suppress high-frequency redraws when output is redirected.&lt;/p&gt;
&lt;h3&gt;19.4 The apparent region lock&lt;/h3&gt;
&lt;p&gt;The 403 responses from the early Node transport caused the API to look geographically restricted. Controlled requests using &lt;code&gt;curl&lt;/code&gt; and .NET disproved that assumption. The actual cause was the earlier HTTP path rather than a regional requirement.&lt;/p&gt;
&lt;p&gt;This was a useful reminder to isolate transport failures before attributing them to application-level restrictions.&lt;/p&gt;
&lt;h2&gt;20. Final architecture&lt;/h2&gt;
&lt;p&gt;The final repository has four logically distinct layers:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;flowchart TD
    A[Live game client] --&gt; B[Asset map]
    A --&gt; C[Client bundle and API metadata]
    B --&gt; D[Catalog]
    D --&gt; E[Downloader]
    C --&gt; F[Authenticated album API]
    F --&gt; G[Card hash store]
    G --&gt; E
    E --&gt; H[Raw mirror]
    H --&gt; I[Text decryption and extraction]
    I --&gt; J[Organized archive]
    I --&gt; K[Readable story scripts]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The important boundary is between acquisition and interpretation. The raw mirror remains close to the upstream layout. All naming, grouping, decryption, and presentation decisions happen later.&lt;/p&gt;
&lt;p&gt;This means a change in extraction rules does not require a new download, and a change in card-hash resolution does not invalidate the rest of the archive.&lt;/p&gt;
&lt;h2&gt;21. What remains unresolved&lt;/h2&gt;
&lt;p&gt;The investigation did not establish a complete source for every hash-bearing resource family in the game. In particular:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;21 card IDs in the v442 asset map were not represented in the 28-character album crawl. They belong to four collaboration characters and Hazuki.&lt;/li&gt;
&lt;li&gt;Six whisper &lt;code&gt;...000.m4a&lt;/code&gt; files were present in the asset map but had no corresponding album entry in the examined data.&lt;/li&gt;
&lt;li&gt;Some specialised hashed systems outside the card and whisper paths were not integrated because they have their own records and were outside the scope of &lt;code&gt;shinymas-dl&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The exact server-side consequences, if any, of calling the album POST repeatedly were not established. The downloader therefore treats the calls as read-oriented but does not assume that the endpoint is formally documented as side-effect free.&lt;/li&gt;
&lt;li&gt;The request codec&apos;s internal SFMT construction was understood far enough to characterize its time-dependent behaviour, but the finished tool deliberately executes the game&apos;s codec instead of maintaining a separate native cryptographic implementation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are boundaries of the investigation rather than claims that the corresponding resources are permanently unavailable.&lt;/p&gt;
&lt;h2&gt;22. Conclusion&lt;/h2&gt;
&lt;p&gt;The difficult part of &lt;code&gt;shinymas-dl&lt;/code&gt; was not writing a parallel downloader. It was identifying the layers between a logical game resource and the object ultimately requested by the browser.&lt;/p&gt;
&lt;p&gt;The first layer was public and straightforward once the client was read carefully: the asset map described 386,102 resources, the filename transformation could be reconstructed from the resource-hashing code, and the text format could be reversed through known compressed output.&lt;/p&gt;
&lt;p&gt;The second layer was more deceptive. The CDN returned 200 for nonexistent paths, making the initial coverage estimate look much better than the actual card-art coverage. Once card requests were compared with their client-side path builders, the missing information was identifiable as a card-level hash.&lt;/p&gt;
&lt;p&gt;The final step was to determine where that metadata entered the browser. The character album exposed hashes for all of its card entries, including entries not owned by the test account. The distinction between metadata availability and UI interaction explained why the network trace appeared to show a condition-gated resource.&lt;/p&gt;
&lt;p&gt;The final implementation reflects that understanding. &lt;code&gt;shinymas-dl&lt;/code&gt; can maintain an ordinary raw mirror, augment it with authenticated card metadata when available, download the corresponding protected resources, and organize the result offline without coupling the archive layout to the game&apos;s current UI.&lt;/p&gt;
&lt;p&gt;As of asset-map v442, the completed album crawl recorded 1,440 hashes for 1,461 card IDs and established a complete hash mapping for every unowned card represented by the 28-character album set. The remaining gaps are specific and identifiable rather than an undifferentiated collection of HTTP 404s.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Appendix A. Selected resource-path examples&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;# Logical resource path
images/content/idols/card/1040010010.jpg

# Ordinary CDN lookup
/assets/&amp;#x3C;sha256-derived-name&gt;?v=&amp;#x3C;version&gt;

# Card-hash lookup after album resolution
images/content/idols/card/&amp;#x3C;card-hash&gt;_1040010010.jpg
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For media resources, the hashed CDN object keeps its media extension, for example &lt;code&gt;.mp4&lt;/code&gt; or &lt;code&gt;.m4a&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;Appendix B. Relevant client-side fields and modules&lt;/h2&gt;
&lt;p&gt;The exact webpack module IDs are build-specific and will change when the client is rebuilt. The following names were useful during this investigation:&lt;/p&gt;
&lt;p&gt;| Client element | Role |
|---|---|
| &lt;code&gt;env.js&lt;/code&gt; | API root, game ID and environment values |
| asset-map loader | Manifest and versioned asset chunks |
| &lt;code&gt;resource_hash&lt;/code&gt; | Ordinary CDN filename hashing and text resource processing |
| &lt;code&gt;request_hash&lt;/code&gt; | API request and response transport encoding |
| character album module | &lt;code&gt;characterAlbums/characters/{characterId}&lt;/code&gt; |
| &lt;code&gt;album/top&lt;/code&gt; | Character list for the album UI |
| &lt;code&gt;idolWhisperVoices&lt;/code&gt; | Whisper-voice records and release conditions |&lt;/p&gt;
&lt;h2&gt;Appendix C. Reproducibility notes&lt;/h2&gt;
&lt;p&gt;All numerical measurements in this document refer to the 27 September 2026 investigation and, where applicable, asset-map v442. The authenticated tests used a dedicated test account. Authentication material itself is intentionally absent from this document.&lt;/p&gt;
&lt;p&gt;The raw session transcript used to prepare this report is archived separately. The report is written from the observed sequence of experiments, corrections, implementation changes, and validation results in that transcript, rather than from assumptions about how the game ought to work.&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Reverse Engineering RcloneView Before Trusting It With My S3 Keys</title><link>https://pure.vermilion10.dev/blog/reverse-engineering-rcloneview-before-trusting-it-with-s3-keys</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/reverse-engineering-rcloneview-before-trusting-it-with-s3-keys</guid><description>A closed-source rclone client wants my S3 credentials. Its developer says nothing is transmitted. I pulled the APK apart and captured its traffic to find out whether that&apos;s true.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;
&lt;p&gt;[!NOTE]
Everything here is against apps I installed on my own device (Galaxy S24 FE,
Android 16), for the purpose of deciding whether to trust them with my own
credentials. No servers were touched that weren&apos;t mine or the vendor&apos;s own
public endpoints. The fake AWS keys used in the traffic test are the
well-known documentation examples.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In the last post, on cloning apps into Samsung&apos;s dual-app profile, I mentioned
almost in passing that I&apos;d gone looking for a trustworthy rclone/S3 client for
Android and that &quot;closed-source options were an easy no.&quot; That was a
lazy heuristic and I knew it while writing it. &quot;Closed source&quot; is not the same
as &quot;untrustworthy,&quot; it just means the cost of finding out is higher. So I decided
to actually pay that cost.&lt;/p&gt;
&lt;p&gt;The app is &lt;a href=&quot;https://play.google.com/store/apps/details?id=com.bdrive.rcloneviewmobile&quot;&gt;RcloneView Mobile&lt;/a&gt;,
a proprietary rclone GUI that had just landed on Google Play. It looked genuinely
good. It also wanted the one thing I&apos;m least willing to hand to a black box: my
S3 access key and secret.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;The claims on the record&lt;/h2&gt;
&lt;p&gt;Before touching a disassembler, it&apos;s worth knowing what you&apos;re testing. Someone
had already asked the obvious question on the vendor&apos;s own forum. In a thread
bluntly titled &lt;a href=&quot;https://forum.rcloneview.com/t/safety-closed-source/405&quot;&gt;&quot;Safety, closed source&quot;&lt;/a&gt;,
a user named Stan_Manioc asked:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Why should we feel safe since you are closed source and you could be leaking
credentials and private cloud files?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After a week of silence he pushed again, saying he couldn&apos;t understand how to
safely hand &quot;config password, cloud api info, to a black box.&quot; On 4 January 2026,
the RcloneView account replied with four specific, falsifiable claims:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&quot;RcloneView itself is only a GUI layer.&quot;&lt;/li&gt;
&lt;li&gt;&quot;All actual operations, authentication, encryption, data transfer, and file
handling, are performed entirely by the rclone binary.&quot;&lt;/li&gt;
&lt;li&gt;&quot;Credentials and config data are stored locally using the OS&apos;s secure storage
mechanisms (such as Keychain on macOS or equivalent).&quot;&lt;/li&gt;
&lt;li&gt;&quot;We do not transmit your credentials or file data to any external servers, nor
do we have visibility into your cloud contents.&quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Their marketing says the same thing in softer words: the
&lt;a href=&quot;https://rcloneview.com/support/blog/cloud-storage-security-checklist-rcloneview&quot;&gt;security checklist post&lt;/a&gt;
describes a &quot;local-first&quot; design where &quot;data flows directly between your machine
and your clouds, no intermediary&quot; and &quot;no third-party server touches your data.&quot;&lt;/p&gt;
&lt;p&gt;That&apos;s a good set of claims because every one of them is checkable. And crucially,
that forum thread is about the &lt;strong&gt;desktop&lt;/strong&gt; app. The mobile app is new, version
1.0.0, and nobody had checked it at all.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Getting the app apart&lt;/h2&gt;
&lt;p&gt;The app was already installed, so no sketchy APK mirrors were needed:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm path com.bdrive.rcloneviewmobile
package:/data/app/~~yOrvTTLcKbYUtCx9JNxnsQ==/...base.apk
package:/data/app/~~.../split_config.arm64_v8a.apk
package:/data/app/~~.../split_config.en.apk
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Six splits, pulled with &lt;code&gt;adb pull&lt;/code&gt;. The arm64 split is 99 MB, which is the first
interesting fact. Nothing about a file browser needs 99 MB of native code unless
it&apos;s carrying an entire Go program.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ unzip -l split_config.arm64_v8a.apk | sort -rn
 70058008   lib/arm64-v8a/libgojni.so
 11316480   lib/arm64-v8a/libflutter.so
 10879920   lib/arm64-v8a/libapp.so
  1717000   lib/arm64-v8a/libsqlite3.so
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So it&apos;s three layers:&lt;/p&gt;
&lt;p&gt;| Layer | What it is |
|---|---|
| &lt;code&gt;classes.dex&lt;/code&gt; | Thin Flutter plugin glue, wrapped in &lt;strong&gt;PairIP&lt;/strong&gt; (Google Play&apos;s anti-tamper) |
| &lt;code&gt;libapp.so&lt;/code&gt; | Dart AOT snapshot, all the app logic |
| &lt;code&gt;libgojni.so&lt;/code&gt; | A real rclone, compiled to a shared library via &lt;code&gt;gomobile&lt;/code&gt; |&lt;/p&gt;
&lt;p&gt;The dex is nearly worthless for analysis, it&apos;s plugin registration and PairIP
scaffolding. The interesting parts are the two blobs that decompilers don&apos;t
touch.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!TIP]
Flutter release builds are AOT-compiled Dart, so &lt;code&gt;jadx&lt;/code&gt; gives you almost
nothing. But the AOT snapshot &lt;strong&gt;retains library and symbol names&lt;/strong&gt;. A plain
strings pass over &lt;code&gt;libapp.so&lt;/code&gt; is far more productive than any decompiler.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That turned out to be the whole ballgame:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ python strings.py libapp.so 6 | grep -oE &quot;package:rcloneview_mobile/[a-z_/]+\.dart&quot; | sort -u | wc -l
246
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;246 source file paths, giving me the app&apos;s entire architecture for free:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;api/secure_storage_service.dart
api/drive_credentials_extractor.dart
api/remote_config_injector.dart
api/license_client.dart
api/device_guid.dart
pairing/api/qr_crypto.dart
util/diag_log.dart
...
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;Where the credentials actually live&lt;/h2&gt;
&lt;p&gt;Two findings settle claim 1 and claim 2 together.&lt;/p&gt;
&lt;p&gt;First, the JNI bridge. The Go layer exports 113 symbols under a package called
&lt;code&gt;ndwrap&lt;/code&gt;, and the config-related ones are exactly what you&apos;d hope:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;ndwrap.SetRemoteConfig
ndwrap.SetRemoteConfigObscured
ndwrap.GetRemoteConfig
ndwrap.inMemoryStorage
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That last symbol is the important one. Combined with &lt;code&gt;RemoteConfigInjector&lt;/code&gt; and a
&lt;code&gt;rclone-inject&lt;/code&gt; tag on the Dart side, it says the rclone configuration is
&lt;strong&gt;injected into the engine in memory&lt;/strong&gt;. There is no &lt;code&gt;rclone.conf&lt;/code&gt; written to disk
at all. I confirmed the negative case by searching the device:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell &quot;find /sdcard -iname &apos;*rclone*&apos; -o -iname &apos;*bdrive*&apos;&quot;
(nothing)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Second, persistence. The app stores state through &lt;code&gt;flutter_secure_storage&lt;/code&gt; under
these keys, all recovered from the snapshot:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;rcloneview.remotes.ini
rcloneview.license.info
bdrive.auth.token
bdrive.session.key
bdrive.device.guid
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And both of that plugin&apos;s Android backends are present in the dex: Tink-backed
&lt;code&gt;EncryptedSharedPreferences&lt;/code&gt; (AES-256-GCM), plus the legacy keystore cipher:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-java&quot;&gt;public final Cipher p() {
    return Cipher.getInstance(&quot;RSA/ECB/OAEPPadding&quot;, &quot;AndroidKeyStoreBCWorkaround&quot;);
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Either path wraps the data with a key held in the AndroidKeyStore, which on an
S24 FE is hardware-backed and non-exportable. So claim 3 holds: credentials are
ciphertext at rest, unlocked by a key that can&apos;t leave the device.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!WARNING]
One caveat worth understanding. &lt;code&gt;SetRemoteConfigObscured&lt;/code&gt; maps to rclone&apos;s
&lt;code&gt;obscure&lt;/code&gt; function, and rclone&apos;s &quot;obscure&quot; is reversible AES with a hardcoded
key. It is obfuscation, not encryption. That&apos;s fine &lt;em&gt;here&lt;/em&gt;, because the
obscured blob is itself inside Keystore-backed storage, but the obscuring layer
contributes exactly zero real protection on its own.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;What it sends home&lt;/h2&gt;
&lt;p&gt;Static analysis found only two first-party domains in the entire Dart snapshot:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;https://rcloneview.com
https://license.rcloneview.com
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Everything else was a cloud provider you&apos;d expect (&lt;code&gt;graph.microsoft.com&lt;/code&gt;,
&lt;code&gt;api.dropboxapi.com&lt;/code&gt;, &lt;code&gt;s3.amazonaws.com&lt;/code&gt;). More importantly, I went looking for
telemetry SDKs and found &lt;strong&gt;none&lt;/strong&gt;: no Firebase Analytics, no Crashlytics, no
AppsFlyer, Adjust, Sentry, Amplitude, or Mixpanel. The &lt;code&gt;firebase-*.properties&lt;/code&gt;
files in the APK are transitive ML Kit dependencies pulled in by the QR scanner,
not analytics.&lt;/p&gt;
&lt;p&gt;That&apos;s a good static result. But static analysis can only prove what code
&lt;em&gt;exists&lt;/em&gt;, never what it &lt;em&gt;does at runtime&lt;/em&gt;. Claim 4 needed a real test.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Actually watching the traffic&lt;/h2&gt;
&lt;p&gt;Here&apos;s where I hit the wall I expected. The app&apos;s network security config
declares no user CA trust anchors, so a normal MITM proxy can&apos;t decrypt anything,
and PairIP exists specifically to make patching the APK unpleasant. I initially
wrote this off as unverifiable.&lt;/p&gt;
&lt;p&gt;That was wrong, and the fix is that &lt;strong&gt;I didn&apos;t need decryption&lt;/strong&gt;. I only needed to
know &lt;em&gt;who it talks to&lt;/em&gt;, and TLS SNI is sent in the clear. PCAPdroid captures
per-app traffic through a local VPN with no root at all:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell am start -n com.emanuelef.remote_capture/.activities.CaptureCtrl \
    -e action start -e app_filter com.bdrive.rcloneviewmobile
$ adb shell ip addr show tun0
87: tun0: &amp;#x3C;POINTOPOINT,UP,LOWER_UP&gt; mtu 10000 ...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;With the capture live and filtered to just that one package, I force-stopped
RcloneView, cold-started it, let it idle, then added an S3 remote using the
canonical fake AWS documentation keys and hit save and verify. It returned
&quot;Access denied,&quot; which is precisely right, that&apos;s a genuine HTTP 403 coming back
from AWS.&lt;/p&gt;
&lt;p&gt;Every connection the app made, for the entire session:&lt;/p&gt;
&lt;p&gt;| Time | Destination | SNI | Sent | Rcvd |
|---|---|---|---|---|
| 15:04:08 | 52.219.37.6 | &lt;code&gt;s3.ap-southeast-1.amazonaws.com&lt;/code&gt; | 2,735 B | 7,230 B |
| 15:04:14 | 52.219.37.6 | &lt;code&gt;s3.ap-southeast-1.amazonaws.com&lt;/code&gt; | 2,815 B | 7,270 B |
| 15:05:43 | 3.5.147.222 | &lt;code&gt;s3.ap-southeast-1.amazonaws.com&lt;/code&gt; | 2,775 B | 7,310 B |
| 15:05:53 | 3.5.147.222 | &lt;code&gt;s3.ap-southeast-1.amazonaws.com&lt;/code&gt; | 2,695 B | 7,230 B |&lt;/p&gt;
&lt;p&gt;Four connections. All to AWS S3. &lt;strong&gt;Zero&lt;/strong&gt; contact with &lt;code&gt;license.rcloneview.com&lt;/code&gt;
or &lt;code&gt;rcloneview.com&lt;/code&gt; across a cold start, three minutes of idle UI use, credential
entry, and browsing.&lt;/p&gt;
&lt;p&gt;The byte counts are worth a glance too: roughly 2.7 KB outbound is about right for
a TLS handshake plus a SigV4-signed S3 request. There&apos;s no room in there for bulk
exfiltration, and nowhere for it to go regardless, since AWS was the only peer.&lt;/p&gt;
&lt;h3&gt;The trap I nearly fell into&lt;/h3&gt;
&lt;p&gt;I run AdGuard as a Private DNS resolver, and I very nearly published the above as
a clean bill of health without thinking it through. The problem: &lt;strong&gt;if AdGuard had
been blocking &lt;code&gt;license.rcloneview.com&lt;/code&gt;, the app would never have opened a TCP
connection to it, and my capture would look identical to a clean one.&lt;/strong&gt; I&apos;d be
reporting &quot;no telemetry&quot; when the real finding was &quot;my DNS filter hid the
telemetry.&quot;&lt;/p&gt;
&lt;p&gt;So I tested the resolver directly rather than trusting the absence:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ curl -s -H &apos;accept: application/dns-json&apos; \
    &quot;https://dns.adguard.com/resolve?name=license.rcloneview.com&amp;#x26;type=A&quot;
{&quot;Answer&quot;:[{&quot;data&quot;:&quot;lb-bdweb-230694326.us-east-1.elb.amazonaws.com.&quot;},
           {&quot;data&quot;:&quot;100.59.159.229&quot;},{&quot;data&quot;:&quot;32.194.61.192&quot;}],&quot;Status&quot;:0}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;AdGuard resolves it, identically to Cloudflare. Nothing was being filtered into
invisibility. The silence in the capture was real silence.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT]
This is the part I&apos;d emphasise to anyone doing this kind of testing. An absence
of evidence in a packet capture is only meaningful if you&apos;ve independently
proven that the evidence &lt;em&gt;could&lt;/em&gt; have appeared. DNS filtering, VPN routing, and
app filters all silently manufacture clean results.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;The verdict on RcloneView&lt;/h2&gt;
&lt;p&gt;All four of the developer&apos;s claims held up under testing. Credentials go straight
to AWS with no vendor server in the path, they&apos;re stored under a hardware-backed
key, rclone does the actual work, and the app did not phone home once.&lt;/p&gt;
&lt;p&gt;That&apos;s a better result than I expected, and I want to state it plainly because my
starting position was prejudicial. Still, the findings weren&apos;t uniformly good:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;allowBackup=&quot;true&quot;&lt;/code&gt;&lt;/strong&gt; with no &lt;code&gt;dataExtractionRules&lt;/code&gt; and no backup agent. App
data is eligible for cloud and device-to-device backup. The practical impact is
limited, since Keystore keys never leave the device so restored ciphertext is
undecryptable, but for an app holding credentials this should be off.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;MANAGE_EXTERNAL_STORAGE&lt;/code&gt;&lt;/strong&gt;, all-files access. Defensible for a file manager,
but it&apos;s the broadest permission in the manifest.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No &lt;code&gt;ACCESS_LOCAL_NETWORK&lt;/code&gt;.&lt;/strong&gt; On Android 17+ that permission gates local
network access, so SMB to a NAS will break on upgrade until they ship a fix.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;rclone 1.73.4&lt;/strong&gt;, a few releases behind.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Developer paths left in the binary:&lt;/strong&gt; &lt;code&gt;/Users/tsjeong/go&lt;/code&gt;,
&lt;code&gt;/Users/tsjeong/workspace&lt;/code&gt;. Cosmetic, but it means no &lt;code&gt;-trimpath&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Version 1.0.0&lt;/strong&gt;, and the forum&apos;s
&lt;a href=&quot;https://forum.rcloneview.com/t/rclone-view-android-app-issues/602&quot;&gt;Android bug thread&lt;/a&gt;
from July 2026 has a user reporting 250 of 1,300 uploads silently failing, no
bulk retry, and no transfer history. Staff acknowledged all of it. The security
posture is fine; the maturity isn&apos;t there yet.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And the structural point that no single test can fix: &lt;strong&gt;today&apos;s clean capture
says nothing about version 1.1.0.&lt;/strong&gt; With a closed-source app you&apos;re not
verifying once, you&apos;re re-verifying forever.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;The open-source comparison&lt;/h2&gt;
&lt;p&gt;For contrast I did the same review on &lt;a href=&quot;https://github.com/chenxiaolong/RSAF&quot;&gt;RSAF&lt;/a&gt;,
where &quot;review&quot; means actually reading the code.&lt;/p&gt;
&lt;p&gt;::github{repo=&quot;chenxiaolong/RSAF&quot;}&lt;/p&gt;
&lt;p&gt;Its credential handling is a straightforwardly stronger design. &lt;code&gt;rclone.conf&lt;/code&gt;
lives in app-private storage and is encrypted with &lt;strong&gt;rclone&apos;s own native config
encryption&lt;/strong&gt;, not merely obscured. The password for that is a randomly generated
128-byte value, which is itself wrapped by Tink AES-256-GCM under an
AndroidKeyStore master key:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-kotlin&quot;&gt;private val hardwareWrappedPassword: Password
    get() {
        ...
        passwordStore.password = RandomUtils.generatePassword(128)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Two real layers instead of one. There are also small signs of a careful author
throughout, like the password type refusing to render itself into a log line:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-kotlin&quot;&gt;data class Password(val value: String) {
    override fun toString(): String = &quot;&amp;#x3C;password&gt;&quot;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The backup handling is the standout, though. Rather than a bare manifest flag,
RSAF ships a &lt;code&gt;BackupAgentHelper&lt;/code&gt; that &lt;strong&gt;refuses to back up at all&lt;/strong&gt; unless the
transport is verifiably encrypted:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-kotlin&quot;&gt;if (data.transportFlags and FLAG_CLIENT_SIDE_ENCRYPTION_ENABLED != 0) {
    Log.i(TAG, &quot;Client-side encrypted backup&quot;)
} else if (data.transportFlags and FLAG_DEVICE_TO_DEVICE_TRANSFER != 0) {
    Log.i(TAG, &quot;Device-to-device transfer&quot;)
} else {
    Log.e(TAG, &quot;Plain-text backup is not allowed&quot;)
    return
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It&apos;s also gated behind a preference that&apos;s off by default. The README is honest
about why: if you &lt;em&gt;do&lt;/em&gt; enable it, the config leaves the device in plaintext and
you&apos;re trusting the backup transport, and the author notes that &quot;some older OEM
Android builds have backup systems that lie about end-to-end encryption.&quot;&lt;/p&gt;
&lt;p&gt;Telemetry: none. The only URLs anywhere in the app source are documentation links
in comments.&lt;/p&gt;
&lt;p&gt;I verified the shipped 4.12 build matches its claims. It carries &lt;strong&gt;rclone 1.75.1&lt;/strong&gt;,
it leaks no developer paths (so &lt;code&gt;-trimpath&lt;/code&gt; is working), it embeds an
&lt;code&gt;assets/archive.tar&lt;/code&gt; snapshot of its own source, and the signature checks out
against the digest published in the README:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ apksigner verify --print-certs RSAF-4.12-arm64-v8a-release.apk
V2 Signer: certificate DN: CN=Andrew Gunnerson, OU=RSAF
V2 Signer: certificate SHA-256 digest:
  b2506499bea1c5a6e658f07be6773fe486999dc124204c6522af7407503ac9f9
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Why my build failed&lt;/h3&gt;
&lt;p&gt;I&apos;d tried to build RSAF from source and couldn&apos;t, which is what sent me down the
&quot;just sideload something&quot; path in the first place. The answer is unglamorous:
&lt;strong&gt;upstream only builds on Linux.&lt;/strong&gt; The CI is &lt;code&gt;runs-on: ubuntu-latest&lt;/code&gt;, full stop.&lt;/p&gt;
&lt;p&gt;The Windows failure is specific. RSAF puts a shim &lt;code&gt;go.exe&lt;/code&gt; first on &lt;code&gt;PATH&lt;/code&gt;
(&lt;a href=&quot;https://github.com/chenxiaolong/RSAF/blob/master/rcbridge/gowrapper/go.go&quot;&gt;&lt;code&gt;rcbridge/gowrapper/go.go&lt;/code&gt;&lt;/a&gt;)
so it can strip absolute paths out of &lt;code&gt;go list -json&lt;/code&gt; output for reproducible
builds. &lt;code&gt;gomobile&lt;/code&gt;&apos;s newer tool-directive check doesn&apos;t survive that shim, and you
get a misleading error:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;gomobile bind requires golang.org/x/mobile in the current module,
but it is not in the module dependency graph.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That message sends you straight to editing &lt;code&gt;go.mod&lt;/code&gt;, which is the wrong move. The
module is fine, as running the query directly proves:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ go list -m golang.org/x/mobile
golang.org/x/mobile v0.0.0-20260821190718-4776eadac327
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Build it in a clean Linux clone and the problem evaporates. Don&apos;t copy a Windows
working tree across either, that trips Gradle&apos;s committed dependency verification
metadata.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Side by side&lt;/h2&gt;
&lt;p&gt;| | RcloneView 1.0.0 | RSAF 4.12 |
|---|---|---|
| Source | Proprietary, PairIP-hardened | GPL-3.0, readable |
| Config at rest | Obscured INI in Keystore storage | Encrypted &lt;code&gt;rclone.conf&lt;/code&gt; + 128-byte Keystore-wrapped password |
| Backup | &lt;code&gt;allowBackup=true&lt;/code&gt;, no rules | Agent refuses plaintext, off by default |
| Telemetry | License server exists, never fired in testing | None |
| rclone | 1.73.4 | 1.75.1 |
| Android 17 ready | Missing &lt;code&gt;ACCESS_LOCAL_NETWORK&lt;/code&gt; | Declared |
| Build hygiene | Dev paths present | &lt;code&gt;-trimpath&lt;/code&gt;, embedded source |
| Verifiable | Only by re-testing each release | Signature + reproducible build |&lt;/p&gt;
&lt;p&gt;One honest asymmetry: these aren&apos;t drop-in equivalents. RSAF is a Storage Access
Framework provider, it exposes remotes to &lt;em&gt;other&lt;/em&gt; apps rather than giving you a
browser UI. RcloneView is a full file manager with photo backup and a rather
clever encrypted QR pairing flow for importing remotes from desktop (PBKDF2 plus
AES-GCM, chunked across multiple QR codes, PIN-authenticated, entirely offline).
If you want the app itself to be the interface, they don&apos;t compete.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;What I&apos;d actually tell you&lt;/h2&gt;
&lt;p&gt;I set out to catch a closed-source app doing something wrong and it didn&apos;t. The
developer&apos;s public claims were accurate, at least for the version I tested, and
the &quot;closed source is an easy no&quot; instinct I started with was intellectually
lazy. Verification beats vibes.&lt;/p&gt;
&lt;p&gt;But I&apos;m still putting my keys in RSAF, for a reason that has nothing to do with
RcloneView&apos;s behaviour: the cost of &lt;em&gt;staying&lt;/em&gt; confident is completely different.
Auditing RcloneView means repeating this entire exercise on every release. Reading
RSAF means reading a diff.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!CAUTION]
None of this replaces bounding your blast radius. Use a dedicated IAM user
scoped to a single bucket, never account-wide or root keys, and rotate them.
Then the worst case is bounded no matter which app you pick, and no amount of
reverse engineering is a substitute for that.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;What I could not verify&lt;/h3&gt;
&lt;p&gt;To be explicit about the limits, since a security writeup without them is
marketing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Payload contents.&lt;/strong&gt; No CA trust means no decryption. I know &lt;em&gt;who&lt;/em&gt; the app
talked to, not what it said. The byte counts constrain the possibilities but
don&apos;t eliminate them.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;That the license server is never contacted.&lt;/strong&gt; &lt;code&gt;LicenseClient&lt;/code&gt; and
&lt;code&gt;DeviceGuidGenerator&lt;/code&gt; exist in the binary. A six-minute window proves they
don&apos;t fire at startup or during ordinary S3 use, nothing more.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;On-disk ciphertext.&lt;/strong&gt; The device isn&apos;t rooted and the app isn&apos;t debuggable, so
I inferred the storage format from the code rather than reading
&lt;code&gt;shared_prefs&lt;/code&gt; directly.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2&gt;Tooling&lt;/h2&gt;
&lt;p&gt;Everything here used free tools and an unrooted phone:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/skylot/jadx&quot;&gt;&lt;code&gt;jadx&lt;/code&gt;&lt;/a&gt; for the dex and manifest&lt;/li&gt;
&lt;li&gt;A 20-line Python strings extractor, since Git Bash ships without &lt;code&gt;strings&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;adb&lt;/code&gt; for pulling splits, permission state and &lt;code&gt;dumpsys&lt;/code&gt; flags&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/emanuele-f/PCAPdroid&quot;&gt;PCAPdroid&lt;/a&gt; for rootless per-app capture&lt;/li&gt;
&lt;li&gt;&lt;code&gt;apksigner&lt;/code&gt; from the Android SDK build-tools&lt;/li&gt;
&lt;li&gt;DNS-over-HTTPS queries to sanity-check the resolver&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The highest-leverage trick by far was realising that Flutter AOT snapshots keep
their symbol names. That one &lt;code&gt;grep&lt;/code&gt; turned an opaque 10 MB blob into a file
listing of the entire application.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Sources: &lt;a href=&quot;https://forum.rcloneview.com/t/safety-closed-source/405&quot;&gt;Safety, closed source&lt;/a&gt; ·
&lt;a href=&quot;https://forum.rcloneview.com/t/rclone-view-android-app-issues/602&quot;&gt;RcloneView Android App issues&lt;/a&gt; ·
&lt;a href=&quot;https://forum.rcloneview.com/t/encrypt-config-file/440&quot;&gt;Encrypt config file?&lt;/a&gt; ·
&lt;a href=&quot;https://rcloneview.com/support/blog/cloud-storage-security-checklist-rcloneview&quot;&gt;Cloud Storage Security Checklist&lt;/a&gt; ·
&lt;a href=&quot;https://github.com/chenxiaolong/RSAF&quot;&gt;RSAF&lt;/a&gt; ·
&lt;a href=&quot;https://play.google.com/store/apps/details?id=com.bdrive.rcloneviewmobile&quot;&gt;Google Play listing&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Cloning Any App Into Samsung&apos;s Dual App Profile</title><link>https://pure.vermilion10.dev/blog/cloning-any-app-into-samsungs-dual-app-profile</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/cloning-any-app-into-samsungs-dual-app-profile</guid><description>Samsung&apos;s Dual Messenger only offers to clone WhatsApp and Facebook. The underlying profile is generic, one ADB command clones anything into it, and one more stops Samsung from undoing it on reboot.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;blockquote&gt;
&lt;p&gt;[!NOTE]
Package names, device model, and Android version in this post are real (Galaxy
S24 FE, One UI 8.5 on Android 16). I&apos;ve redacted local network addresses since
they aren&apos;t relevant to the technique&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;at first this had nothing to do with dual apps at all. It started with trying to
share a couple of Windows folders with my phone over SMB, tightened down to that
one device instead of the whole Wi-Fi network. That went fine, but it left me
wanting proper cloud storage access from Android instead of ad-hoc file shares, so
I went looking for a trustworthy rclone/S3 client. Closed-source options were an
easy no, so I picked two open-source ones, &lt;a href=&quot;https://github.com/chenxiaolong/RSAF&quot;&gt;RSAF&lt;/a&gt;
and a maintained fork of Round-Sync, and decided to build both from source rather
than sideload someone else&apos;s binary.&lt;/p&gt;
&lt;p&gt;RSAF never built, its toolchain is pinned to an Android SDK version so new that
even the maintainer&apos;s own CI has to pin the emulator to a lower API &quot;until the
action supports it.&quot; Round-Sync did build, after fixing a missing &lt;code&gt;GOOS=android&lt;/code&gt;
in its dependency resolution and a 16KB page-alignment issue in its native
library. None of that is this post. This post is about a side effect I noticed
while reinstalling Round-Sync over and over during that debugging: two identical
app icons&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Two icons, one APK&lt;/h2&gt;
&lt;p&gt;I&apos;d run &lt;code&gt;adb install -r app-debug.apk&lt;/code&gt;, and the launcher would show a duplicate,
same name, same icon, one with a small badge overlay. My first guess was Secure
App, Samsung&apos;s Knox-isolated profile. Nah, it was wrong, the badge didn&apos;t look like
Secure Folder&apos;s lock icon, and a permission error from &lt;code&gt;pm list packages&lt;/code&gt; alone
doesn&apos;t distinguish between Samsung&apos;s various profile types anyway. Also, the Secure App only apear in Secure Folder, not on home screen. So I checked
properly:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm list users
Users:
	UserInfo{0:null:4c13} running
	UserInfo{95:DUAL_APP:20001010} running
	UserInfo{150:Secure Folder:61010} running
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Three profiles. &lt;code&gt;150&lt;/code&gt; is Secure Folder, confirmed separately. &lt;code&gt;95&lt;/code&gt; is named
&lt;code&gt;DUAL_APP&lt;/code&gt;, and a package check confirmed it:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm list packages --user 95 | grep felixnuesse
package:de.felixnuesse.extract.debug
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;the app showed up on both the main profile and this dual-app profile, though I&apos;d
never opened Samsung&apos;s dual-app settings for it&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT]
&lt;code&gt;Settings &gt; Advanced features &gt; Dual Messenger&lt;/code&gt; on this device lists exactly two
apps: WhatsApp and Facebook. There is no toggle for anything else. And yet an
unrelated file-manager app was sitting in that profile.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;Ruling out the obvious explanation&lt;/h2&gt;
&lt;p&gt;My first guess was that Dual Messenger secretly supports more than it lets on, or
maybe some Good Lock module got switched on somewhere along the way. Neither one
held up though:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Dual Messenger app list has exactly those two apps, no hidden &quot;add more&quot;
screen.&lt;/li&gt;
&lt;li&gt;Apps installed the normal way, tapping a downloaded APK through the system
installer, never get cloned. A local test with an unrelated open-source app
confirmed this: install by tapping the file, no duplicate appears.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The only apps that ended up duplicated were the ones I had installed via
&lt;code&gt;adb install&lt;/code&gt; in the same session. That&apos;s the actual variable. The dual-app
setting had nothing to do with it.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell dumpsys user
  UserInfo{95:xxx:20001010} serialNo=95 isPrimary=false parentId=0
    Type: android.os.usertype.profile.CLONE
    Flags: 536875024 (DUALAPP_PROFILE|INITIALIZED|PROFILE)
    Created: +537d20h20m49s210ms ago
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;android.os.usertype.profile.CLONE&lt;/code&gt; is AOSP&apos;s generic clone-profile primitive.
Samsung didn&apos;t invent it. Dual Messenger is one front-end that writes into it,
with a curated app allowlist baked into its own UI. The profile itself has no
such allowlist, it&apos;s a plain Android user, and &lt;code&gt;pm&lt;/code&gt; doesn&apos;t care which front-end
you&apos;d normally use to populate it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;The actual technique&lt;/h2&gt;
&lt;p&gt;If the profile is generic, installing into it directly should work without going
anywhere near Dual Messenger&apos;s UI at all. Android&apos;s package manager has a command
built for exactly this: cloning an app that&apos;s &lt;em&gt;already installed&lt;/em&gt; into another
user, without touching its APK.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm list packages --user 0 | grep discord
package:com.discord

$ adb shell pm install-existing --user 95 com.discord
Package com.discord installed for user: 95

$ adb shell pm list packages --user 95 | grep discord
package:com.discord
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That&apos;s the whole trick: no APK re-download, no re-signing. It&apos;s the same
installed package already verified and running on your main profile, registered
against a second user, so there&apos;s no risk of ending up with a different build.
The new icon appears in the launcher, with its own isolated &lt;code&gt;/data/user/95/&lt;/code&gt;
storage: a second, independent login.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;flowchart LR
    subgraph Device
        M[Main profile&amp;#x3C;br/&gt;user 0]
        D[Dual App profile&amp;#x3C;br/&gt;user 95, CLONE type]
        S[Secure Folder&amp;#x3C;br/&gt;user 150]
    end
    UI[Dual Messenger UI] --&gt;|writes only&amp;#x3C;br/&gt;WhatsApp/Facebook| D
    ADB[&quot;pm install-existing --user 95 &amp;#x26;lt;pkg&amp;#x26;gt;&quot;] --&gt;|writes any package| D
    style ADB fill:#22c55e,color:#000
    style UI fill:#6b7280,color:#fff
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The Dual Messenger UI and the raw &lt;code&gt;pm&lt;/code&gt; command both write into the same profile.
One gates it behind an allowlist, the other doesn&apos;t gate it at all.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!TIP]
If the app isn&apos;t installed on the main profile yet, &lt;code&gt;pm install-existing&lt;/code&gt; has
nothing to clone. Install it normally first (Play Store or a verified APK), then
run the command above.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;Then it disappeared&lt;/h2&gt;
&lt;p&gt;I wrote the first draft of this post, logged into the cloned Discord, and moved
on. Some time later the second icon was gone. The app and its data were gone
too:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell dumpsys package com.discord | grep &quot;User 95&quot;
    User 95: ceDataInode=-1 deDataInode=-1 installed=false ...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;ceDataInode=-1&lt;/code&gt; means the profile&apos;s copy &lt;em&gt;and its data&lt;/em&gt; were deleted. Same for
the two other apps I&apos;d cloned. Profile 95 was back to exactly one third-party
package: WhatsApp.&lt;/p&gt;
&lt;p&gt;Android&apos;s log buffer only holds a few hours, so whatever did it had scrolled
away. I re-cloned Discord and started watching. Nothing happened for fourteen
minutes of normal use. Then I rebooted the phone, and two minutes after boot:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;18:13:19.385 E DualAppManagerService: getAllWhitelistedPackages : empty
18:13:19.385 E SemDualAppManager: getAllWhitelistedPackages : null returned. Return default
18:13:19.395 I PackageManager: START INSTALL-EXISTING PACKAGE: userId{95} pkg{com.samsung.android.permissioncontroller}
18:13:19.395 I PackageManager:           Request from pId{12519 / com.samsung.android.da.daagent}
18:13:19.447 I PackageManager: START DELETE PACKAGE: observer{138464497}
18:13:19.447 I PackageManager: pkg{com.discord}, user{95}, caller{1000} flags{0}
18:13:19.451 I ActivityManager: Force stopping com.discord appid=10378 user=95: deletePackageX
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That&apos;s the whole mechanism in seven lines. &lt;code&gt;com.samsung.android.da.daagent&lt;/code&gt;,
the Dual Messenger agent, runs a reconciliation pass on every boot:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Ask the system service for the allowlist. It comes back empty, so it falls
back to the built-in default (WhatsApp and Facebook).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;install-existing&lt;/code&gt; its own helper packages into the profile, using the exact
same command I&apos;d used, which is a nice confirmation that this is the intended
way to populate it.&lt;/li&gt;
&lt;li&gt;Delete every package in the profile that isn&apos;t on the list. Data included.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The profile is generic, the restriction lives in a separate app, and that
separate app is also a janitor that shows up after every reboot. Nothing in the
fourteen minutes of runtime triggered it, only the boot did.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;sequenceDiagram
    participant Boot
    participant DA as daagent
    participant PM as PackageManager
    participant P95 as Profile 95
    Boot-&gt;&gt;DA: start
    DA-&gt;&gt;PM: getAllWhitelistedPackages()
    PM--&gt;&gt;DA: empty → default [WhatsApp, Facebook]
    DA-&gt;&gt;PM: install-existing helpers → user 95
    DA-&gt;&gt;PM: for pkg in profile 95 not in list: deletePackage(pkg, 95)
    PM-&gt;&gt;P95: remove com.discord + data
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I looked for somewhere to write to that allowlist. &lt;code&gt;dumpsys dual_app&lt;/code&gt; prints
nothing, there are no &lt;code&gt;settings&lt;/code&gt; keys for it, and daagent&apos;s data directory is
off-limits without root. Samsung either baked the list into the APK or feeds it
from their servers. Either way, it isn&apos;t ours to edit.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Making it stick&lt;/h2&gt;
&lt;p&gt;If the allowlist can&apos;t be changed, the other lever is the thing that enforces
it. daagent is an ordinary app, and ordinary apps can be disabled per user:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm install-existing --user 95 com.discord
Package com.discord installed for user: 95

$ adb shell pm disable-user --user 0 com.samsung.android.da.daagent
Package com.samsung.android.da.daagent new state: disabled-user
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then reboot. Two minutes after boot, no &lt;code&gt;START DELETE PACKAGE&lt;/code&gt;, no daagent
process, and both clones still there and working, including the WhatsApp one
that Samsung set up itself.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT]
Disabling daagent doesn&apos;t uninstall anything or touch the profile&apos;s data, it
only stops the app that &lt;em&gt;would&lt;/em&gt; delete things. In the &quot;will my WhatsApp clone
get wiped&quot; sense, this makes the profile safer, not riskier. The one thing
that &lt;em&gt;can&lt;/em&gt; still wipe it is turning Dual Messenger off in Settings, which
removes the whole profile, and that&apos;s true with or without this trick.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Disabling daagent costs you a few things (from what i researched on the internet): the Dual Messenger settings page
probably won&apos;t work, and clone notifications might be less reliable since
daagent handles some of the profile&apos;s plumbing. The badge overlay on clone icons
may go away too. All of it comes back with the mirror command:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm enable com.samsung.android.da.daagent
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;which also means the next reboot will sweep Discord out again.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!TIP]
Back up the WhatsApp clone (Settings → Chats → Chat backup, from &lt;em&gt;inside the
clone&lt;/em&gt;) before doing any of this. The commands above don&apos;t threaten it. A
clone profile with no backup is one accidental Settings toggle away from gone
regardless.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If you&apos;d rather not touch a system app at all, the fallback is to re-run
&lt;code&gt;install-existing&lt;/code&gt; after each reboot and log in again. Safe, but annoying.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;What&apos;s still a guess&lt;/h2&gt;
&lt;p&gt;One thing I never nailed down is why plain &lt;code&gt;adb install&lt;/code&gt; (no &lt;code&gt;--user&lt;/code&gt; flag) put
apps into the dual-app profile in the first place, which is what started all of
this. The working theory is that &lt;code&gt;adb install&lt;/code&gt; on this OEM build targets all
users while the tap-to-install path targets only the current one, but I don&apos;t
have a controlled test to call that proven. Everything in the two sections
above, the reboot sweep and the fix for it, is verified from the logs and
reproducible on demand.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Caveats before you do this&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Permission prompts don&apos;t stick.&lt;/strong&gt; The cloned Discord asked for photo access
every time and never got it. &lt;code&gt;dumpsys package&lt;/code&gt; showed why: every runtime
permission for user 95 was &lt;code&gt;granted=false&lt;/code&gt; with no &lt;code&gt;USER_SET&lt;/code&gt; flag, so my
&quot;Allow&quot; taps were never recorded, and &lt;code&gt;POST_NOTIFICATIONS&lt;/code&gt; was denied the same
way. Grant them from the shell instead:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;$ adb shell pm grant --user 95 com.discord android.permission.READ_MEDIA_IMAGES
$ adb shell pm grant --user 95 com.discord android.permission.POST_NOTIFICATIONS
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Same for &lt;code&gt;READ_MEDIA_VIDEO&lt;/code&gt;, &lt;code&gt;CAMERA&lt;/code&gt;, &lt;code&gt;RECORD_AUDIO&lt;/code&gt;, whatever the app asks
for. Force-close the clone afterwards.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Notifications can still be less reliable&lt;/strong&gt; on a clone-profile instance of an
app that was never designed with Samsung&apos;s dual-app framework in mind, even
once the permission is granted. Log in and send yourself a test message
before relying on it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;This isn&apos;t an officially supported configuration&lt;/strong&gt; for arbitrary apps, only
Dual Messenger&apos;s own allowlisted apps get Samsung&apos;s testing and support.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Check the app&apos;s terms of service&lt;/strong&gt; if the clone is for running a second
account. Some services restrict multi-accounting, and this post is about the
Android mechanism, not a statement on any particular app&apos;s policy.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Removing a clone is the mirror command:
&lt;code&gt;adb shell pm uninstall --user 95 &amp;#x3C;package&gt;&lt;/code&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2&gt;Takeaways&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;A restrictive front-end UI doesn&apos;t mean the underlying platform mechanism is
restrictive too. Samsung&apos;s Dual Messenger allowlist lives in its own app. The
clone-profile primitive it writes to has none.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;pm install-existing --user &amp;#x3C;id&gt; &amp;#x3C;package&gt;&lt;/code&gt; clones an already-installed app
into any user profile without touching the APK. Samsung&apos;s own agent uses the
same command.&lt;/li&gt;
&lt;li&gt;If something you set up reverts on its own, suspect a boot-time job before
suspecting a bug. &lt;code&gt;logcat&lt;/code&gt; right after a reboot, filtered on the package
name, found the culprit in seconds.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;pm disable-user&lt;/code&gt; on the enforcing app is a legitimate, reversible fix when
the policy it enforces isn&apos;t editable. Know what else that app does before
you switch it off.&lt;/li&gt;
&lt;/ol&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/><atom:updated>2026-09-17T00:00:00.000Z</atom:updated></item><item><title>Building a Local-First Gacha Tracker: Reverse-Engineering Wuthering Waves&apos; Convene Log</title><link>https://pure.vermilion10.dev/blog/building-a-local-first-gacha-tracker</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/building-a-local-first-gacha-tracker</guid><description>Why grep finds nothing in Wuthering Waves&apos; Client.log, how to decode it in one line, and the archive bug that nearly ate six months of history.</description><pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;::github{repo=&quot;vermilion10/wuwa-pulltrack&quot;}&lt;/p&gt;
&lt;h1&gt;Building a Local-First Gacha Tracker&lt;/h1&gt;
&lt;p&gt;Every gacha tracker asks you to paste a URL into someone else&apos;s website. That URL
contains your account identifier and a signed token. The site fetches your pull
history, stores it, and shows you a dashboard.&lt;/p&gt;
&lt;p&gt;Nothing about that pipeline &lt;em&gt;requires&lt;/em&gt; a third party. The game already writes
everything needed onto your own disk. This post is the record of removing the
middleman, including the two places where my first assumption was wrong, and the
bug that would have silently destroyed the very data the tool exists to preserve.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!NOTE]
All account identifiers and tokens in this post are redacted. The &lt;code&gt;record_id&lt;/code&gt; in
particular is a short-lived credential, treat it like a password.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;The premise&lt;/h2&gt;
&lt;p&gt;The hypothesis was simple:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The game writes a gacha-history URL somewhere local.&lt;/li&gt;
&lt;li&gt;Tracker sites just read that URL and call the game publisher&apos;s own API.&lt;/li&gt;
&lt;li&gt;Therefore the whole loop can run locally.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;All three turned out to be true. Getting to them took longer than expected.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Method: three dead ends&lt;/h2&gt;
&lt;h3&gt;Dead end 1, grep finds nothing&lt;/h3&gt;
&lt;p&gt;The community scripts all describe the same approach: search &lt;code&gt;Client.log&lt;/code&gt; for a
URL matching &lt;code&gt;aki-gm-resources.../aki/gacha/index.html#/record&lt;/code&gt;. So I searched
every log in the game&apos;s log directory.&lt;/p&gt;
&lt;p&gt;Zero matches. Not even the bare substring &lt;code&gt;gacha&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The file was clearly alive, modified seconds earlier, 3.5 MB. So the data was
there and my reader was wrong. Dumping the first bytes explained why:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;EF BB BF B4 97 95 97 93 8B 95 9D 8B 97 DE C2 DE DA ...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A UTF-8 BOM (&lt;code&gt;EF BB BF&lt;/code&gt;), then bytes that are &lt;strong&gt;not valid UTF-8&lt;/strong&gt;. The log is
obfuscated. A rotated backup, closed, not being written, showed the same shape,
so this wasn&apos;t an artifact of reading a file mid-write.&lt;/p&gt;
&lt;h3&gt;Dead end 2, the webview&apos;s storage&lt;/h3&gt;
&lt;p&gt;The convene history screen is a webview, so its storage seemed promising.
&lt;code&gt;Client/Saved/LocalStorage/LocalStorage.db&lt;/code&gt; is a plain SQLite file:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-sql&quot;&gt;CREATE TABLE LocalStorage(key text primary key not null, value text not null);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It holds ~250 keys, graphics settings, red-dot state, &lt;code&gt;RecentlyLoginUID&lt;/code&gt;, and
teasingly &lt;code&gt;GachaPoolOpenRecord_&amp;#x3C;uid&gt;&lt;/code&gt;. But that last one is just a list of banner
IDs you&apos;ve &lt;em&gt;viewed&lt;/em&gt;. No URL, no token.&lt;/p&gt;
&lt;h3&gt;Dead end 3, the SDK&apos;s debug log&lt;/h3&gt;
&lt;p&gt;Older scripts mention a second path:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Client/Binaries/Win64/ThirdParty/KrPcSdk_Global/KRSDKRes/KRSDKWebView/debug.log
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This one &lt;em&gt;is&lt;/em&gt; plaintext, a Chromium debug log. It still exists, but contains only
GPU errors and &lt;code&gt;spdy_session&lt;/code&gt; warnings. Whatever used to log the record URL there
no longer does.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT]
Three plausible sources, all dry. At this point the only remaining lead was the
obfuscated log itself.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;Finding: the cipher is one line&lt;/h2&gt;
&lt;p&gt;A single-byte XOR brute force scored no key above ~0.83 printable ratio, so not a
trivial cipher. But there was a much better lever available than brute force:
&lt;strong&gt;known plaintext&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;UE4 logs open with their own timestamp, and the rotated backup file was &lt;em&gt;named&lt;/em&gt;
&lt;code&gt;Client-backup-2026.08.21-15.21.53-...&lt;/code&gt;. So the plaintext almost certainly began:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;[2026.08.21-15.21.53:...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Aligning that against the cipher bytes:&lt;/p&gt;
&lt;p&gt;| Position | Cipher | Plain | &lt;code&gt;cipher ^ plain&lt;/code&gt; |
|---|---|---|---|
| 0 | &lt;code&gt;B4&lt;/code&gt; | &lt;code&gt;[&lt;/code&gt; (&lt;code&gt;5B&lt;/code&gt;) | &lt;code&gt;EF&lt;/code&gt; |
| 1 | &lt;code&gt;97&lt;/code&gt; | &lt;code&gt;2&lt;/code&gt; (&lt;code&gt;32&lt;/code&gt;) | &lt;code&gt;A5&lt;/code&gt; |
| 2 | &lt;code&gt;95&lt;/code&gt; | &lt;code&gt;0&lt;/code&gt; (&lt;code&gt;30&lt;/code&gt;) | &lt;code&gt;A5&lt;/code&gt; |
| 3 | &lt;code&gt;97&lt;/code&gt; | &lt;code&gt;2&lt;/code&gt; (&lt;code&gt;32&lt;/code&gt;) | &lt;code&gt;A5&lt;/code&gt; |
| 4 | &lt;code&gt;93&lt;/code&gt; | &lt;code&gt;6&lt;/code&gt; (&lt;code&gt;36&lt;/code&gt;) | &lt;code&gt;A5&lt;/code&gt; |
| 5 | &lt;code&gt;8B&lt;/code&gt; | &lt;code&gt;.&lt;/code&gt; (&lt;code&gt;2E&lt;/code&gt;) | &lt;code&gt;A5&lt;/code&gt; |
| 10 | &lt;code&gt;DE&lt;/code&gt; | &lt;code&gt;1&lt;/code&gt; (&lt;code&gt;31&lt;/code&gt;) | &lt;code&gt;EF&lt;/code&gt; |
| 11 | &lt;code&gt;C2&lt;/code&gt; | &lt;code&gt;-&lt;/code&gt; (&lt;code&gt;2D&lt;/code&gt;) | &lt;code&gt;EF&lt;/code&gt; |&lt;/p&gt;
&lt;p&gt;Two keys, &lt;code&gt;0xA5&lt;/code&gt; and &lt;code&gt;0xEF&lt;/code&gt;, but alternating on no positional schedule. Sorting by
which key applied reveals the rule instantly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Key &lt;code&gt;0xA5&lt;/code&gt; applies to: &lt;code&gt;32 30 36 2E 38 3A&lt;/code&gt;, all &lt;strong&gt;even&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Key &lt;code&gt;0xEF&lt;/code&gt; applies to: &lt;code&gt;5B 31 2D 35 33 5D&lt;/code&gt;, all &lt;strong&gt;odd&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The key is selected by the &lt;em&gt;plaintext&lt;/em&gt; byte&apos;s parity. That sounds circular, you
need the plaintext to pick the key, except both keys have bit 0 set, so XOR always
flips the low bit. &lt;strong&gt;The cipher byte&apos;s own low bit tells you which key was used.&lt;/strong&gt;&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;def decode(raw: bytes) -&gt; str:
    body = raw[3:] if raw[:3] == b&quot;\xef\xbb\xbf&quot; else raw   # BOM is not obfuscated
    return bytes((b ^ 0xA5) if (b &amp;#x26; 1) else (b ^ 0xEF)
                 for b in body).decode(&quot;utf-8&quot;)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Result: valid UTF-8, 84,002 lines, readable UE4 log, including Chinese-language
entries, which is why the naive &quot;printable ASCII ratio&quot; heuristic capped at 0.86
and made the decode look wrong at first glance.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!TIP]
When a brute force stalls, look for known plaintext. A log file that embeds its
own timestamp in its filename is a gift.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And there it was:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;https://aki-gm-resources-oversea.aki-game.net/aki/gacha/index.html#/record
  ?player_id=&amp;#x3C;REDACTED&gt;
  &amp;#x26;record_id=&amp;#x3C;REDACTED&gt;
  &amp;#x26;svr_id=&amp;#x3C;REDACTED&gt;
  &amp;#x26;resources_id=&amp;#x3C;REDACTED&gt;
  &amp;#x26;gacha_type=1&amp;#x26;lang=en&amp;#x26;platform=PC
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;The API&lt;/h2&gt;
&lt;p&gt;Those four parameters are everything the record endpoint needs:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;POST https://gmserver-api.aki-game2.net/gacha/record/query
{
  &quot;playerId&quot;:     &quot;&amp;#x3C;player_id&gt;&quot;,
  &quot;cardPoolId&quot;:   &quot;&amp;#x3C;resources_id&gt;&quot;,
  &quot;cardPoolType&quot;: 1,               # 1..8, the actual banner selector
  &quot;languageCode&quot;: &quot;en&quot;,
  &quot;recordId&quot;:     &quot;&amp;#x3C;record_id&gt;&quot;,
  &quot;serverId&quot;:     &quot;&amp;#x3C;svr_id&gt;&quot;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Probing all eight pool types showed &lt;code&gt;cardPoolType&lt;/code&gt; is the real selector,
&lt;code&gt;cardPoolId&lt;/code&gt; doesn&apos;t need to vary with it. Each call returns that banner&apos;s
&lt;strong&gt;complete window&lt;/strong&gt;, newest first:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;cardPoolType&quot;: &quot;Resonators Accurate Modulation&quot;,
  &quot;resourceId&quot;: 21030013,
  &quot;qualityLevel&quot;: 3,
  &quot;resourceType&quot;: &quot;Weapon&quot;,
  &quot;name&quot;: &quot;Pistols of Night&quot;,
  &quot;count&quot;: 1,
  &quot;time&quot;: &quot;2026-08-20 10:23:17&quot;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;The one irreducible constraint&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;record_id&lt;/code&gt; is minted when you open Convene History in-game, and expires in roughly
a day. &lt;strong&gt;No tool avoids this&lt;/strong&gt;, not this one, not the hosted trackers. It is a
property of the API, not of running locally.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;sequenceDiagram
    participant P as Player
    participant G as Game client
    participant L as Client.log
    participant T as Tracker (local)
    participant K as Kuro API

    P-&gt;&gt;G: Open Convene History
    G-&gt;&gt;L: write signed URL (obfuscated)
    P-&gt;&gt;T: run `sync`
    T-&gt;&gt;L: read + XOR-decode
    L--&gt;&gt;T: player_id, record_id, svr_id
    loop pool types 1..8
        T-&gt;&gt;K: POST /gacha/record/query
        K--&gt;&gt;T: full window for that banner
    end
    T-&gt;&gt;T: archive into SQLite
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;Architecture&lt;/h2&gt;
&lt;p&gt;No server, no client, no backend. One CLI that produces a report.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;flowchart LR
    A[Client.log] --&gt;|XOR decode| B[convene URL]
    B --&gt;|parse params| C[Kuro record API]
    C --&gt;|~6 month window| D[(pulls.db&amp;#x3C;br/&gt;append-only)]
    D --&gt; E[report.html]
    style A fill:#4a3aa7,color:#fff
    style D fill:#eda100,color:#000
    style E fill:#1baf7a,color:#fff
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A static-file pipeline beats a local web server here: nothing to keep running,
nothing listening on a port, and the output is a file you can archive or email.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;The bug that mattered&lt;/h2&gt;
&lt;p&gt;Kuro&apos;s API serves a &lt;strong&gt;rolling window of roughly six months&lt;/strong&gt;. Older pulls stop
coming back. That single fact is the tool&apos;s entire reason to exist, and it broke
my first schema.&lt;/p&gt;
&lt;p&gt;The API gives &lt;strong&gt;no pull ID&lt;/strong&gt;. So identity has to be derived. My first attempt keyed
each row on its position from the oldest record:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;pull_index = total - i   # WRONG
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The reasoning: history only grows at the newest end, so position-from-oldest is
stable. That reasoning holds &lt;em&gt;only while the window doesn&apos;t slide&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;When the oldest records age out, every index shifts by the number dropped. The
upsert then writes new content onto old keys. Replaying that rule against a
simulated slide on a real 550-row banner:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-text&quot;&gt;v1 keys written on full window : 550
v1 keys written on slid window : 350
keys whose CONTENT changed     : 350   &amp;#x3C;-- silent overwrites
  e.g. pull_index 350: was &apos;Originite: Type I&apos; @ 2026-06-13
                       now &apos;Pistols of Night&apos;  @ 2026-08-20
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;350 of 550 rows silently corrupted&lt;/strong&gt;, precisely the old history the archive
exists to protect. No error, no constraint violation. Just wrong data.&lt;/p&gt;
&lt;h3&gt;The fix: content-addressed identity&lt;/h3&gt;
&lt;p&gt;Anchor identity to &lt;strong&gt;absolute time&lt;/strong&gt;, never to position:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;PRIMARY KEY (player_id, pool_type, time, seq)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;seq&lt;/code&gt; is needed because timestamps alone don&apos;t identify a pull:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A 10-pull shares &lt;strong&gt;one&lt;/strong&gt; timestamp across all ten records.&lt;/li&gt;
&lt;li&gt;The same &lt;code&gt;resource_id&lt;/code&gt; can appear &lt;strong&gt;three times within that single second&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So within each timestamp, pulls are numbered &lt;code&gt;1..n&lt;/code&gt; oldest-first. That key survives
any window movement, because a timestamp never changes.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;stateDiagram-v2
    [*] --&gt; Fetched: API returns window
    Fetched --&gt; Existing: key already in archive
    Fetched --&gt; New: key unseen
    Existing --&gt; Kept: refresh last_seen only
    New --&gt; Stored: insert with first_seen
    Kept --&gt; [*]
    Stored --&gt; [*]
    note right of Kept
        Rows the API stopped
        returning are never
        touched. That is the
        archive.
    end note
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Verified against the real database:&lt;/p&gt;
&lt;p&gt;| Scenario | Expected | Result |
|---|---|---|
| Drop oldest 200 from response | archive unchanged | 550 rows, byte-identical |
| Slid window + 3 new pulls | old kept, new added | 553 rows, old intact |
| Re-sync identical data | no-op | 0 new rows |&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[!WARNING]
If you build one of these: &lt;strong&gt;never key gacha rows by position in the API
response.&lt;/strong&gt; It works perfectly until the day it destroys your history, and it
fails silently when it does.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;Findings from 755 pulls&lt;/h2&gt;
&lt;p&gt;With the archive in place, the data is worth a look. Across three banners:&lt;/p&gt;
&lt;p&gt;| Banner | Pulls | 5★ | Rate | Avg pity |
|---|---|---|---|---|
| Featured Resonator | 550 | 10 | 1.82% | 54.1 |
| Featured Weapon | 84 | 1 | 1.19% | 63.0 |
| Standard | 121 | 2 | 1.65% | 45.5 |
| &lt;strong&gt;All&lt;/strong&gt; | &lt;strong&gt;755&lt;/strong&gt; | &lt;strong&gt;13&lt;/strong&gt; | &lt;strong&gt;1.72%&lt;/strong&gt; | &lt;strong&gt;53.5&lt;/strong&gt; |&lt;/p&gt;
&lt;p&gt;The observed rate $\hat{p} = 13/755 = 1.72%$ sits close to the advertised 0.8%
base rate once pity is accounted for. The Wilson 95% interval for this sample is
wide:&lt;/p&gt;
&lt;p&gt;$$
\hat{p} \pm z\sqrt{\frac{\hat{p}(1-\hat{p})}{n}} = 0.0172 \pm 1.96\sqrt{\frac{0.0172 \times 0.9828}{755}}
$$&lt;/p&gt;
&lt;p&gt;which gives roughly $[0.86%,\ 2.65%]$, with 13 successes, this sample cannot
distinguish much. &lt;strong&gt;Thirteen five-stars is an anecdote, not a measurement.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The pity distribution is more interesting than the rate:&lt;/p&gt;
&lt;p&gt;| Pity bin | 11–20 | 31–40 | 51–60 | 61–70 | 71–80 |
|---|---|---|---|---|---|
| Count | 3 | 1 | 1 | 5 | 3 |&lt;/p&gt;
&lt;p&gt;Eight of thirteen landed at pity 61 or higher, and the 21–30, 41–50 bins are empty.
That bimodal shape, a few early, a cluster late, a gap between, is the visible
fingerprint of a &lt;strong&gt;soft pity&lt;/strong&gt; mechanic: the rate stays low, then ramps sharply
somewhere in the 60s. You can see the mechanic in the shape even when the sample is
far too small to pin down the rate.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;What it looks like&lt;/h2&gt;
&lt;p&gt;The output is a single self-contained &lt;code&gt;report.html&lt;/code&gt;, no CDN, no server, no
network. Per banner: current pity against the 80-pull hard cap, 5★ rate, average
pity, and every 5★ as a bar whose length is the pity it landed at. Plus an
all-banners summary with monthly volume and the histogram above.&lt;/p&gt;
&lt;p&gt;Dark mode is a selected palette rather than an inverted one, and the two data
colors (gold for 5★, violet for 4★) were validated for colorblind separation
rather than eyeballed, worst-pair ΔE 41.0 in light, 27.3 in dark.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Takeaways&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Obfuscation is not encryption.&lt;/strong&gt; A parity-selected two-key XOR is one line to
undo once you have known plaintext. The filename was the known plaintext.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prefer the boring architecture.&lt;/strong&gt; A CLI that writes a static HTML file needs
no server, no port, no process manager, and no deploy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interrogate your primary key against time.&lt;/strong&gt; &quot;History only grows at the end&quot;
was true on the day I wrote it and false six months later. Any key derived from
&lt;em&gt;position in a paginated response&lt;/em&gt; is a bug waiting for a rolling window.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Local-first means the archive outlives the API.&lt;/strong&gt; The upstream keeps six
months. The SQLite file keeps everything.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The whole thing is ~600 lines of Python with no dependencies outside the standard
library. The interesting part was never the code, it was the twenty minutes of
staring at hex.&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Automated Daily Check-in System for Arknights Endfield: Architecture and Implementation</title><link>https://pure.vermilion10.dev/blog/article-akefcheckin</link><guid isPermaLink="true">https://pure.vermilion10.dev/blog/article-akefcheckin</guid><description>Technical article documenting the akef-checkin system: OAuth 2.0 flow, HMAC-SHA256 signing, GitHub Actions deployment, and multi-account automation for Arknights Endfield daily rewards.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;::github{repo=&quot;vermilion10/akef-checkin&quot;}&lt;/p&gt;
&lt;h2&gt;Abstract&lt;/h2&gt;
&lt;p&gt;This article presents the design and implementation of &lt;code&gt;akef-checkin&lt;/code&gt;, an automated daily check-in system for &lt;em&gt;Arknights Endfield&lt;/em&gt; (SKport platform). The system employs a full OAuth 2.0 authorization flow with HMAC-SHA256 request signing to authenticate against Hypergryph and SKport APIs, eliminating the need for manual token refresh. Deployed as a serverless workflow via GitHub Actions, the system executes daily at 16:00 UTC with multi-account support and Discord webhook notifications. We detail the cryptographic signing mechanism, the four-step authentication pipeline, and the deployment architecture.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;1 Introduction&lt;/h2&gt;
&lt;p&gt;Daily check-in systems for live-service games typically require authenticated API requests to claim rewards. The &lt;em&gt;Arknights Endfield&lt;/em&gt; platform (operated by SKport/Hypergryph)[^1][^2] implements a multi-layer authentication scheme involving:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Long-lived account token&lt;/strong&gt; (&lt;code&gt;ACCOUNT_TOKEN&lt;/code&gt;) — stored by user, valid for weeks to months&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Short-lived OAuth code&lt;/strong&gt; — obtained via &lt;code&gt;POST /user/oauth2/v2/grant&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Session credential&lt;/strong&gt; (&lt;code&gt;cred&lt;/code&gt;) — generated via &lt;code&gt;POST /user/auth/generate_cred_by_code&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Signing token&lt;/strong&gt; (&lt;code&gt;signToken&lt;/code&gt;) — retrieved via &lt;code&gt;GET /auth/refresh&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Game role identifier&lt;/strong&gt; — resolved via &lt;code&gt;GET /game/player/binding&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Traditional approaches store short-lived tokens, requiring frequent manual updates. Our system regenerates all transient credentials on every execution using only the long-lived &lt;code&gt;ACCOUNT_TOKEN&lt;/code&gt;, achieving zero-maintenance operation[^3].&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;2 System Architecture&lt;/h2&gt;
&lt;h3&gt;2.1 High-Level Overview&lt;/h3&gt;
&lt;p&gt;The system comprises four modular components:&lt;/p&gt;
&lt;p&gt;| Component | File | Responsibility |
|-----------|------|----------------|
| Entry Point | &lt;code&gt;src/index.ts&lt;/code&gt; | Orchestration, multi-account iteration, result aggregation |
| Authentication | &lt;code&gt;src/auth.ts&lt;/code&gt; | Four-step OAuth flow, cryptographic signing |
| Notification | &lt;code&gt;src/discord.ts&lt;/code&gt; | Discord webhook delivery |
| Configuration | &lt;code&gt;src/constants.ts&lt;/code&gt; | API endpoints, headers, platform constants |&lt;/p&gt;
&lt;h3&gt;2.2 Execution Pipeline&lt;/h3&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;flowchart TD
    A[Start: Read SKPORT_TOKENS] --&gt; B{For each account}
    B --&gt; C[Step 1: Get OAuth Code]
    C --&gt; D[Step 2: Generate Cred]
    D --&gt; E[Step 3: Get Sign Token]
    E --&gt; F[Step 4: Get Game Role]
    F --&gt; G[Sign Attendance Request]
    G --&gt; H[POST /game/endfield/attendance]
    H --&gt; I{Response Code}
    I --&gt;|0| J[Success: Parse Rewards]
    I --&gt;|1001, 10001| K[Already Signed In]
    I --&gt;|10002| L[Token Expired]
    I --&gt;|Other| M[Failure: Log Error]
    J --&gt; N[Aggregate Result]
    K --&gt; N
    L --&gt; N
    M --&gt; N
    N --&gt; O{More Accounts?}
    O --&gt;|Yes| B
    O --&gt;|No| P[Send Discord Webhook]
    P --&gt; Q[Exit]
    
    style A fill:#f97316,color:#fff
    style Q fill:#22c55e,color:#fff
    style J fill:#3b82f6,color:#fff
    style K fill:#3b82f6,color:#fff
    style L fill:#f97316,color:#fff
    style M fill:#ef4444,color:#fff
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;3 Authentication Flow&lt;/h2&gt;
&lt;h3&gt;3.1 Sequence Diagram&lt;/h3&gt;
&lt;p&gt;The complete authentication sequence involves two identity providers: Hypergryph (&lt;code&gt;as.gryphline.com&lt;/code&gt;) for OAuth grant and SKport (&lt;code&gt;zonai.skport.com&lt;/code&gt;) for game-specific credentials.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-mermaid&quot;&gt;sequenceDiagram
    participant Script
    participant Gryphline as as.gryphline.com
    participant SKport as zonai.skport.com

    Script-&gt;&gt;Gryphline: POST /user/oauth2/v2/grant&amp;#x3C;br/&gt;{token: ACCOUNT_TOKEN, appCode, type: 0}
    Gryphline--&gt;&gt;Script: {code: &quot;oauth_code&quot;} (one-time use)

    Script-&gt;&gt;SKport: POST /web/v1/user/auth/generate_cred_by_code&amp;#x3C;br/&gt;{code: oauth_code, kind: 1}
    SKport--&gt;&gt;Script: {cred: &quot;session_credential&quot;, token: &quot;...&quot;}

    Script-&gt;&gt;SKport: GET /web/v1/auth/refresh&amp;#x3C;br/&gt;(cred, timestamp, platform, vName)
    SKport--&gt;&gt;Script: {token: &quot;sign_token&quot;}

    Script-&gt;&gt;SKport: GET /api/v1/game/player/binding&amp;#x3C;br/&gt;(cred + signed request)
    SKport--&gt;&gt;Script: {list: [{appCode: &quot;endfield&quot;, bindingList: [...]}]}

    Script-&gt;&gt;SKport: POST /web/v1/game/endfield/attendance&amp;#x3C;br/&gt;(cred + gameRole + signed request)
    SKport--&gt;&gt;Script: {code: 0, data: {signInCount, awardIds, resourceInfoMap}}
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;3.2 Cryptographic Signing (V2)&lt;/h3&gt;
&lt;p&gt;All signed requests use the &lt;strong&gt;V2 signature algorithm&lt;/strong&gt;: HMAC-SHA256 followed by MD5.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-typescript&quot;&gt;export function generateSignV2(
    path: string,
    timestamp: string,
    platform: string,
    vName: string,
    salt: string
): string {
    const headerJson = JSON.stringify({ platform, timestamp, dId: &quot;&quot;, vName });
    const s = `${path}${timestamp}${headerJson}`;
    const hmac = crypto.createHmac(&quot;sha256&quot;, salt).update(s).digest(&quot;hex&quot;);
    return crypto.createHash(&quot;md5&quot;).update(hmac).digest(&quot;hex&quot;);
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Mathematical formulation:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
\text{sign} = \text{MD5}\big(\text{HMAC-SHA256}(\text{salt}, \text{path} \parallel \text{timestamp} \parallel \text{JSON}({\text{platform}, \text{timestamp}, \text{dId}, \text{vName}}))\big)
$$&lt;/p&gt;
&lt;p&gt;Where:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$\parallel$ denotes string concatenation&lt;/li&gt;
&lt;li&gt;$\text{salt} = \text{signToken}$ (per-session secret)&lt;/li&gt;
&lt;li&gt;$\text{path}$ = API endpoint path (e.g., &lt;code&gt;/web/v1/game/endfield/attendance&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;$\text{timestamp}$ = Unix epoch seconds as string&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;[!NOTE]
The double-hash construction (HMAC-SHA256 → MD5) provides defense-in-depth: even if MD5 collision resistance is compromised, the HMAC-SHA256 layer preserves integrity assuming the salt remains secret.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;3.3 Request Structure&lt;/h3&gt;
&lt;p&gt;Each signed request includes the following headers:&lt;/p&gt;
&lt;p&gt;| Header | Value |
|--------|-------|
| &lt;code&gt;cred&lt;/code&gt; | Session credential from Step 2 |
| &lt;code&gt;platform&lt;/code&gt; | &lt;code&gt;&quot;3&quot;&lt;/code&gt; (constant) |
| &lt;code&gt;vName&lt;/code&gt; | &lt;code&gt;&quot;1.2.0&quot;&lt;/code&gt; (client version) |
| &lt;code&gt;timestamp&lt;/code&gt; | Current Unix timestamp (seconds) |
| &lt;code&gt;sign&lt;/code&gt; | V2 signature (hex) |
| &lt;code&gt;sk-game-role&lt;/code&gt; | Game role identifier (format: &lt;code&gt;3_{roleId}_{serverId}&lt;/code&gt;) |
| &lt;code&gt;sk-language&lt;/code&gt; | Locale code (&lt;code&gt;en&lt;/code&gt;, &lt;code&gt;ja&lt;/code&gt;, &lt;code&gt;zh_Hans&lt;/code&gt;, etc.) |
| &lt;code&gt;dId&lt;/code&gt; | Empty string (device ID placeholder) |&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;4 Implementation Details&lt;/h2&gt;
&lt;h3&gt;4.1 Multi-Account Processing&lt;/h3&gt;
&lt;p&gt;The entry point processes accounts sequentially with a 2-second delay between requests to avoid rate limiting:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-typescript&quot;&gt;for (let i = 0; i &amp;#x3C; accounts.length; i++) {
    const account = accounts[i];
    const result = await performCheckin(account);
    results.push(result);

    if (i &amp;#x3C; accounts.length - 1) {
        await new Promise(resolve =&gt; setTimeout(resolve, 2000));
    }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;4.2 Response Parsing&lt;/h3&gt;
&lt;p&gt;On successful check-in (HTTP 200, &lt;code&gt;code === 0&lt;/code&gt;), the response contains reward information:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-typescript&quot;&gt;interface AttendanceResponse {
    code: number;
    message?: string;
    data?: {
        signInCount?: number;
        awardIds?: Array&amp;#x3C;{ id: string }&gt;;
        resourceInfoMap?: Record&amp;#x3C;string, { name: string; count: number }&gt;;
    };
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Rewards are mapped by &lt;code&gt;awardIds&lt;/code&gt; → &lt;code&gt;resourceInfoMap&lt;/code&gt; and formatted as &lt;code&gt;ItemName xCount&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;4.3 Error Handling&lt;/h3&gt;
&lt;p&gt;The system categorizes responses into four classes:&lt;/p&gt;
&lt;p&gt;| Code / Condition | Category | Action |
|------------------|----------|--------|
| &lt;code&gt;code === 0&lt;/code&gt; | Success | Parse rewards, report day count |
| &lt;code&gt;code ∈ {1001, 10001}&lt;/code&gt; or &quot;already&quot; in message | Already Signed In | Report success (idempotent) |
| &lt;code&gt;code === 10002&lt;/code&gt; | Token Expired | Alert user to update &lt;code&gt;ACCOUNT_TOKEN&lt;/code&gt; |
| Other / Network Error | Failure | Log error, continue to next account |&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;5 Deployment Architecture&lt;/h2&gt;
&lt;h3&gt;5.1 GitHub Actions Workflow&lt;/h3&gt;
&lt;p&gt;The system runs serverlessly via GitHub Actions[^5] on &lt;code&gt;ubuntu-latest&lt;/code&gt; with Bun runtime:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-yaml&quot;&gt;name: Endfield Auto Check-in

on:
  schedule:
    - cron: &apos;0 16 * * *&apos;  # Daily at 16:00 UTC (00:00 WIB)
  workflow_dispatch:

jobs:
  checkin:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: oven-sh/setup-bun@v1
        with:
          bun-version: latest
      - run: bun run src/index.ts
        env:
          SKPORT_TOKENS: ${{ secrets.SKPORT_TOKENS }}
          DISCORD_WEBHOOK_URL: ${{ secrets.DISCORD_WEBHOOK_URL }}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Schedule interpretation:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;0 16 * * *&lt;/code&gt; → Minute 0, Hour 16 (UTC) → &lt;strong&gt;Midnight WIB (UTC+7)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;workflow_dispatch&lt;/code&gt; enables manual triggering via GitHub UI&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;5.2 Secrets Management&lt;/h3&gt;
&lt;p&gt;Two repository secrets are required:&lt;/p&gt;
&lt;p&gt;| Secret | Required | Format |
|--------|----------|--------|
| &lt;code&gt;SKPORT_TOKENS&lt;/code&gt; | Yes | JSON array of account objects |
| &lt;code&gt;DISCORD_WEBHOOK_URL&lt;/code&gt; | No | Discord webhook URL |&lt;/p&gt;
&lt;p&gt;Account object schema:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;accountToken&quot;: &quot;string (from Network tab)&quot;,
  &quot;accountName&quot;: &quot;string (display name)&quot;,
  &quot;language&quot;: &quot;string (optional, default: en)&quot;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;6 Token Acquisition Procedure&lt;/h2&gt;
&lt;p&gt;The &lt;code&gt;ACCOUNT_TOKEN&lt;/code&gt; is &lt;strong&gt;not&lt;/strong&gt; visible in browser cookies. It must be extracted from the OAuth grant request:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Navigate to &lt;code&gt;https://game.skport.com/endfield/sign-in&lt;/code&gt; while logged in&lt;/li&gt;
&lt;li&gt;Open DevTools (F12) → Network tab&lt;/li&gt;
&lt;li&gt;Filter for &lt;code&gt;grant&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Reload page (F5)&lt;/li&gt;
&lt;li&gt;Select &lt;code&gt;POST https://as.gryphline.com/user/oauth2/v2/grant&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Open &lt;strong&gt;Payload&lt;/strong&gt; tab (not Response)&lt;/li&gt;
&lt;li&gt;Copy &lt;code&gt;token&lt;/code&gt; value from JSON body:
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;token&quot;: &quot;ACCOUNT_TOKEN_HERE&quot;,
  &quot;appCode&quot;: &quot;6eb76d4e13aa36e6&quot;,
  &quot;type&quot;: 0
}
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;blockquote&gt;
&lt;p&gt;[!IMPORTANT]
The &lt;code&gt;ACCOUNT_TOKEN&lt;/code&gt; is long-lived (weeks to months). The script regenerates all short-lived credentials (&lt;code&gt;cred&lt;/code&gt;, &lt;code&gt;signToken&lt;/code&gt;, &lt;code&gt;gameRole&lt;/code&gt;) on every run, so token rotation is rarely needed.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2&gt;7 Local Development&lt;/h2&gt;
&lt;h3&gt;7.1 Prerequisites&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://bun.sh/&quot;&gt;Bun&lt;/a&gt; v1.0+&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;7.2 Configuration&lt;/h3&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;cp .env.example .env
# Edit .env with your ACCOUNT_TOKEN
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;7.3 Execution&lt;/h3&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;bun run start
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Expected output:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Starting Arknights Endfield Auto Check-in...

Processing check-in for account: MyAccount...
  [1/4] Getting OAuth code...
  [2/4] Generating cred...
  [3/4] Getting sign token...
  [4/4] Getting player binding (game role)...
  Auth complete! Game role: 3_12345678_9

**Endfield:** Check-in for `MyAccount`
✅ Success!
 Day: 15
 Rewards: Credit x1000, Synthetic Fuel x5

Sending webhook to Discord...
Done!
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;8 Security Analysis&lt;/h2&gt;
&lt;h3&gt;8.1 Threat Model&lt;/h3&gt;
&lt;p&gt;| Asset | Threat | Mitigation |
|-------|--------|------------|
| &lt;code&gt;ACCOUNT_TOKEN&lt;/code&gt; | Exfiltration via logs | Never logged; only used in memory |
| &lt;code&gt;signToken&lt;/code&gt; | Replay attacks | Bound to timestamp; expires quickly |
| &lt;code&gt;cred&lt;/code&gt; | Session hijacking | Short-lived; regenerated per run |
| Discord webhook | Spam/abuse | Optional; rate-limited by Discord |&lt;/p&gt;
&lt;h3&gt;8.2 Cryptographic Properties&lt;/h3&gt;
&lt;p&gt;The V2 signature[^4] provides:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Integrity&lt;/strong&gt;: Any modification to path, timestamp, or headers invalidates signature&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Replay resistance&lt;/strong&gt;: Timestamp inclusion limits window to ~1 second (server-side validation)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key separation&lt;/strong&gt;: &lt;code&gt;signToken&lt;/code&gt; derived per-session, not reusable across accounts&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;8.3 Operational Security&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;No persistent storage of credentials&lt;/li&gt;
&lt;li&gt;All secrets injected via GitHub Actions encrypted secrets&lt;/li&gt;
&lt;li&gt;Network requests use HTTPS with certificate validation (default in &lt;code&gt;fetch&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;User-Agent mimics legitimate browser (Firefox 147)&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2&gt;9 Related Work&lt;/h2&gt;
&lt;p&gt;| Project | Platform | Approach |
|---------|----------|----------|
| &lt;code&gt;akef-checkin&lt;/code&gt; (this work) | Arknights Endfield (SKport) | Full OAuth flow regeneration |
| &lt;code&gt;hoyolab-auto-checkin&lt;/code&gt; | Genshin Impact / Star Rail | Cookie-based with manual refresh |
| &lt;code&gt;arknights-auto-sign&lt;/code&gt; | Arknights (Global/CN) | Token caching with TTL |&lt;/p&gt;
&lt;p&gt;Our approach differs by &lt;strong&gt;eliminating token caching entirely&lt;/strong&gt; — every execution performs a fresh OAuth grant, trading ~3 extra HTTP requests for zero-maintenance operation.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;10 Conclusion&lt;/h2&gt;
&lt;p&gt;The &lt;code&gt;akef-checkin&lt;/code&gt; system demonstrates a practical application of OAuth 2.0 and request signing for game automation. Key contributions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Zero-maintenance authentication&lt;/strong&gt; — single long-lived token generates all ephemeral credentials&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cryptographically sound signing&lt;/strong&gt; — HMAC-SHA256 + MD5 V2 algorithm with per-session keys&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Serverless deployment&lt;/strong&gt; — GitHub Actions cron with Bun runtime, no infrastructure management&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multi-account support&lt;/strong&gt; — sequential processing with rate-limit awareness&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Observability&lt;/strong&gt; — structured logging and Discord notifications&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Future work includes: exponential backoff for transient failures, Prometheus metrics export, and support for additional SKport game titles.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Appendix A: API Endpoint Reference&lt;/h2&gt;
&lt;p&gt;| Constant | URL |
|----------|-----|
| &lt;code&gt;GRANT&lt;/code&gt; | &lt;code&gt;https://as.gryphline.com/user/oauth2/v2/grant&lt;/code&gt; |
| &lt;code&gt;GENERATE_CRED&lt;/code&gt; | &lt;code&gt;https://zonai.skport.com/web/v1/user/auth/generate_cred_by_code&lt;/code&gt; |
| &lt;code&gt;REFRESH_TOKEN&lt;/code&gt; | &lt;code&gt;https://zonai.skport.com/web/v1/auth/refresh&lt;/code&gt; |
| &lt;code&gt;BINDING&lt;/code&gt; | &lt;code&gt;https://zonai.skport.com/api/v1/game/player/binding&lt;/code&gt; |
| &lt;code&gt;ATTENDANCE&lt;/code&gt; | &lt;code&gt;https://zonai.skport.com/web/v1/game/endfield/attendance&lt;/code&gt; |&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Appendix B: Type Definitions&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-typescript&quot;&gt;interface AccountConfig {
    accountToken: string;
    accountName: string;
    language?: string;
}

interface AuthResult {
    cred: string;
    signToken: string;
    gameRole: string;
}

interface ResourceInfo {
    name: string;
    count: number;
}

interface AttendanceData {
    signInCount?: number;
    awardIds?: Array&amp;#x3C;{ id: string }&gt;;
    resourceInfoMap?: Record&amp;#x3C;string, ResourceInfo&gt;;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;p&gt;[^1]: Hypergryph OAuth 2.0 Implementation, &lt;code&gt;as.gryphline.com&lt;/code&gt; (reverse-engineered)
[^2]: SKport Web API v1, &lt;code&gt;zonai.skport.com&lt;/code&gt; (reverse-engineered)
[^3]: RFC 6749 — The OAuth 2.0 Authorization Framework
[^4]: RFC 2104 — HMAC: Keyed-Hashing for Message Authentication
[^5]: GitHub Actions Documentation — Scheduled Events (&lt;code&gt;on.schedule&lt;/code&gt;)&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;This article documents the &lt;code&gt;akef-checkin&lt;/code&gt; repository as of version 1.0.0. Source code available at the project repository.&lt;/em&gt;&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item></channel></rss>