EDGE-FIRST NEURAL VISION • PRODUCTION SPEC

UtilVision Architecture & Benchmarks

A lightweight, two-stage deep learning pipeline engineered to detect and parse utility meter LCD counters locally on edge hardware with zero API dependencies.

100%
Display Recall
78.7ms
Edge GPU Latency
96.5%
Master Exact Match
$0.00
Cloud API Cost
// 01 DOMAIN CHALLENGE

Why Off-the-Shelf OCR Fails on Utility Registers

General-purpose OCR engines achieve less than 40% accuracy on digital utility meters due to three distinct physical failure modes:

7-Segment Disconnection
Unlit segment gaps confuse linguistic OCR dictionaries into hallucinating Latin letters (e.g., classifying 0 into N, or 8 into B).
Standard OCR EngineMisread
O8.B kWh
UtilVision Stage 3100% Match
08.8
Faceplate Clutter & Unit Symbols
Printed serial numbers, barcodes, brand logos, and adjacent measurement units (kWh, kW) distract text detectors and cause catastrophic multiplier errors.
Standard OCR EngineDistracted
SN-984210
UtilVision Stage 2Isolated
LCD BBox Only
Strict Leading Zero Preservation
Utility registers require verbatim leading zero preservation for billing reconciliation. Standard numerical parsers routinely strip leading zeros.
Standard ParserTruncated
7888.54
UtilVision Stage 3Preserved
007888.54
// 02 PIPELINE ARCHITECTURE

Two-Stage Decoupled Inference Pipeline

Interactive Pipeline Transformation Explorer
STEP 01
Handheld Capture
Raw Input Capture
Raw field photo of single-phase residential meter with ambient reflections.
640x640 Input
Input Buffer
STEP 02
Stage 2 Detector
Stage 2 Localization
Anchor-free FPN isolates the active LCD counter, ignoring faceplate logos.
Conf: 86.8%
Detection Score
STEP 03
96px Normalization
Normalized LCD Strip
Fixed 96 px canonical height with dynamic width preserves sub-pixel decimals.
410x123 -> 96px
Dynamic Aspect
STEP 04
Stage 3 Recognizer
Decoded Sequence
007888.54
Confidence: 100.0%
11-token CTC vocabulary enforces strict zero-retention and zero hallucination.
Exact Match
Ledger Verified
Stage 2 Detector (2.62M params / 10.6 MB) Anchor-free FPN with conditional Cascade TTA optical zoom fallback for weathered or distant registers.
Canonical 96px Normalization Standardized 96 px canonical height with dynamic aspect width preserves sub-pixel decimal point separation.
Stage 3 Recognizer (15.51M params / 31.4 MB) SVTR visual token attention with 11-token CTC vocabulary (0 to 9, dot) mathematically prevents hallucination.
// 03 DATASET & PROTOCOL

Tri-Domain Taxonomy & Stratification Protocol

To guarantee real-world generalization across diverse field hardware, the dataset is structured across a tri-domain taxonomy encompassing 4,094 curated images:

Strict Physical Meter-Grouped Stratification Protocol

•
Physical Meter Partitioning: Splits are grouped strictly by unique meter hardware identifier. Images of any single physical meter exist exclusively within train, validation, or the test split.
•
150-Meter Held-Out Corpus: Completely unseen utility equipment ensures unbiased zero-leakage evaluation.
•
Verbatim Ledger Equality: Scored on exact character, decimal point, and leading-zero equality without fuzzy numeric tolerance.
Held-Out Test Corpus:150 Unique Physical Meters
Stratification Rule:Zero Meter ID Overlap
Evaluation Metric:Verbatim Equality
Display Localization:100% Recall (IoU > 0.50)
Digit Recognition:99.4% Normalized Levenshtein
Industrial Set C:96.2% Exact Match
// 04 EVOLUTION & BENCHMARKS

Model Evolution Across 10 Generations

Verbatim exact-match accuracy evaluated consistently on the master benchmark pool across engineering iterations. Tap any generation below to inspect its technical breakthrough:

Recognition Accuracy Across 10 Model Generations
Exact Match (%)
Scroll chart or tap pills to inspect G1 to G10 →
100% 75% 50% 25% 0% 96.5% G1 G2 G3 G4 G5 G6 G7 G8 G9 G10
GENERATION 10 Production SOTA
Champion Stage 3 Recognizer (SOTA)
96.5%
+2.4%
Exact Match
99.4%
Digit Acc
96.2%
Industrial Set C

Core Technical Innovation

Fine-tuned 30 epochs with synthetic specular glare injection and random segment cutout regularization with quantized export.

Physical Field Challenge Solved

Mastered dimmed segments and intense camera flash reflections; achieved 96.2% exact match on industrial three-phase equipment.

10 of 10
View Complete Benchmark Progression Table (10 Generations) ▼
GenArchitecture & Training StrategyExact MatchDigit AccIndustrial (Set C)
1Off-the-shelf CRNN baseline24.1%61.4%8.2%
2Custom 7-segment digit dictionary52.8%81.3%19.5%
3ResNet-18 visual feature extractor68.3%88.9%31.2%
4SVTR visual token interaction neck74.1%91.2%38.6%
5Targeted photometric and blur augmentations81.4%94.1%46.2%
6Dual-domain domestic and commercial balancing85.9%95.8%49.8%
7Industrial high-voltage domain integration89.4%97.1%50.4%
8High-density annotation boundary audit91.8%97.8%51.7%
9Canonical 96px aspect-preserving crop normalization94.1%98.6%84.6%
10Champion Stage 3 Recognizer (synthetic glare & segment cutout)96.5%99.4%96.2%
Open-Access Model Weights on Hugging Face Hub
Stage 2 Detector (YOLO11n) and Stage 3 Recognizer (PP-OCRv6 SVTR-CTC FP16/FP32) are published under AGPL-3.0 copyleft terms.
View on Hugging Face →
// 05 ARCHITECTURE COMPARISON

Edge WebAssembly vs. Cloud Vision Architecture

UtilVision eliminates cloud API subscriptions and cellular connectivity bottlenecks by executing the entire neural pipeline directly within the client browser session. Empirical roundtrip benchmarks to cloud endpoints reveal 768 ms in initial TLS setup alone, yielding 1,100 ms to 2,200 ms broadband turnaround (and 2.5s to 5.0s+ over cellular), alongside outright failure in subterranean utility spaces.

Inference Speed
78.7 ms
15x to 30x faster than cloud API roundtrips (1.1s to 2.2s).
Ongoing Cost
$0.00
Zero per-read fees (vs $1.50 to $2.50+ per 1k on Cloud Vision).
Field Reliability
100% Offline
Guaranteed operation in concrete basements and utility shafts.
Data Privacy
Client-Only
Zero customer meter photos ever leave the local device.
Architecture Metric UtilVision (WebGPU / WASM Fallback) Cloud Vision APIs Python Server OCR
Inference Latency WebGPU: 78.7 ms GPU | WASM: 214 ms CPU 1,100 ms to 2,200 ms (Broadband)
2,500 ms to 5,000+ ms (Cellular)
320 ms + Network Transfer
API Execution Cost $0.00 / Zero Ongoing Cost $1.50 to $2.50+ per 1,000 Reads Server Instance Hosting Fees
Basement & Offline Support 100% Offline (PWA Cache) Fails Without Cellular Link Fails Without Cellular Link
Customer Privacy Zero Image Uploads Transmitted to Remote Cloud Stored on Backend Server
Client Installation Zero Install (Any Web Browser) API Key Integration Required Heavy Native Python Packages
// 06 SPECIFICATIONS

Benchmark Hardware Specifications & Methodology

Standardized edge execution environments and measurement protocols for all reported latencies:

Edge GPU Inference Environment (78.7 ms)

GPU Accelerator:Apple M2 / NVIDIA RTX 4060 Mobile
Runtime Provider:ONNX Runtime Web 1.21 (WebGPU / WASM Hybrid)
Precision:FP16 (Half Precision)
Latency Breakdown:Stage 2: 31.2ms | Stage 3: 47.5ms
Measurement Protocol:performance.now() over 50 warm passes

CPU WebAssembly Baselines (Desktop vs. Mobile)

Desktop Processor:Intel Core i7-13700H / AMD Ryzen 7 7840U
Desktop Provider:WASM SIMD (4 Web Worker Threads)
Desktop Latency:214.3ms (Stage 2: 74ms, Stage 3: 140ms)
Mobile Provider:WASM SIMD (4 Cores via Credentialless COEP)
OnePlus 11R:1.0s to 1.5s (Snapdragon 8+ Gen 1)
OnePlus 13R:Sub-1s / 650ms to 950ms (Snapdragon 8 Gen 3)
Mobile TTA Fallback:1.8s to 2.2s (Weathered / Distant Zoom)

Cache Storage & Memory Footprint

Pipeline Parameters:18.13M Total (Stage 2: 2.62M, Stage 3: 15.51M)
Storage Mechanism:Browser CacheStorage API (Offline PWA)
Stage 2 Detector:10.6 MB (ONNX INT8 / FP16)
Stage 3 Recognizer:31.4 MB (FP16) / 62.3 MB (FP32)
Total Cache Footprint:104.4 MB Verified Storage
Active Tab RAM:~180 MB Active Heap Allocation

Run the pipeline on your own device with live camera or photo input.

Open Live Reader →