A clean, modular Python runtime pipeline and wrapper for the Windows 11 Snipping Tool OCR engine (oneocr.onemodel).
Note: The DLL binaries and model files are Microsoft intellectual property.
This repository contains only the extraction tooling — the binaries are fetched from the official Microsoft Store at runtime and never redistributed.
You can run or install OneOCR directly using uv:
# Run the preparation tool directly from the repository:
uv run --with-editable . oneocr --prepare
# Run OCR directly on an image:
uv run --with-editable . oneocr path/to/screenshot.png --lang ru
# Install globally as a command-line tool:
uv tool install git+https://github.com/bropines/oneocr-onnx-python.git
# Add as a dependency to your uv project:
uv add git+https://github.com/bropines/oneocr-onnx-python.git# Install as a CLI tool and library directly from GitHub:
pip install git+https://github.com/bropines/oneocr-onnx-python.git
# Or install locally in editable mode:
pip install -e .If you have installed OneOCR, you can run the preparation process directly via the CLI:
oneocr --prepare(This queries the Microsoft Store API, downloads the Snipping Tool package, decrypts the ONNX models, and reconstructs the language vocabularies to your user config folder ~/.config/oneocr/ or local models/ folder if running from a cloned workspace).
Alternatively, you can run the local helper script:
python utils/prepare_files.py# Using the installed global CLI command:
oneocr path/to/screenshot.png --lang ru
# Or via python module execution:
python -m oneocr path/to/screenshot.png --lang ruFor detailed documentation on configuration, hardware acceleration, and the full Python class interface, see api.md.
Important
Platform Requirements:
- Preparation & Decryption: The extraction and decryption step (
utils/prepare_files.py) utilizes Windows APIs and native DLLs, meaning it must be run on a Windows machine. - OCR Engine & Inference: Once the sub-models are extracted into the
models/folder, they are standard platform-independent ONNX files. The coreoneocrPython package runs on Windows, macOS, Linux, and Docker using only standard Python libraries andonnxruntime.
Or use the library directly in your own code:
from oneocr import OneOCR
with OneOCR() as ocr:
result = ocr.recognize_file("screenshot.png")
print(result.full_text)
print(f"Image angle: {result.image_angle}°")
for line in result.lines:
print(f"\n[Line] {line.text}")
x, y, w, h = line.bbox.as_rect()
print(f" Position: x={x:.0f} y={y:.0f} w={w:.0f} h={h:.0f}")
for word in line.words:
print(f" '{word.text}' conf={word.confidence:.3f}")OneOCR_Deobfuscated/
│
├── oneocr/ ← Python package (Modular, pure ONNX Runtime)
│ ├── __init__.py ← Package entry point
│ ├── common.py ← Quad, Word, Line, OcrResult classes
│ ├── vocab.py ← Vocabulary loading
│ ├── detector.py ← TextDetector class
│ ├── classifier.py ← ScriptClassifier class
│ ├── recognizer.py ← TextRecognizer class
│ └── corrector.py ← OrientationCorrector class
│
├── docker/ ← Docker container configurations
│ ├── Dockerfile ← Production inference image
│ ├── Dockerfile.prepare ← Model download/decryption image (via Wine)
│ ├── .dockerignore ← Docker ignore patterns
│ └── docker-prepare-entrypoint.sh ← Wine container entrypoint script
│
├── docs/ ← Documentation files
│ ├── api.md ← API Reference and GPU configuration
│ └── WRITEUP.md ← In-depth reverse engineering analysis
│
├── utils/ ← Automated setup and inspection utilities
│ ├── download_and_extract_oneocr.py ← Downloads MS Store package
│ ├── extract_all.py ← Memory patch decryptor
│ ├── inspect_models.py ← Model layout and metadata inspector
│ └── prepare_files.py ← Unified model decryption runner
│
├── tests/ ← Test scripts and samples
│ ├── test_robustness.py
│ └── test_real_image.py ← Verify inference runs properly
│
├── pyproject.toml ← PEP 621 python package configuration
├── agents.md ← Agent pairing rules and guidelines
├── README.md ← This file
│
├── bin/ ← Raw DLL binaries (gitignored)
└── models/ ← Decrypted & organized ONNX sub-models
┌──────────────────────────────────────────┐
│ oneocr Package Pipeline │
│ │
│ 1. text_detector.onnx │
│ FPN detector → bounding quads │
│ │ │
│ 2. script_classifier.onnx │
│ Per-crop: script ID (10 classes) │
│ orientation / flip │
│ │ │
│ 3. recognizer_<lang>.onnx (CRNN+CTC) │
│ Input: [1, 3, 60, W] │
│ Output: logsoftmax over vocab │
│ │ │
│ 4. CTC Decode & Word alignment (Python) │
└────────────────────┬─────────────────────┘
│
Lines + Words + BBox
| Model | File | Vocab | Role |
|---|---|---|---|
| Detector | text_detector.onnx |
— | Finds text regions via FPN |
| Classifier | script_classifier.onnx |
10 | Script type + orientation |
| CJK | recognizer_cjk.onnx |
32 632 | Chinese / Japanese / Korean |
| Cyrillic | recognizer_cyrillic.onnx |
548 | Cyrillic + extended |
| Latin+ | recognizer_latin.onnx |
415 | Extended Latin block |
| Others | recognizer_*.onnx |
179–244 | Smaller script families |
Main engine class. Auto-detects models/ or ~/.config/oneocr/.
engine = OneOCR() # auto-detect
engine = OneOCR(default_language="ru") # force Russian (bypasses classifier)
engine = OneOCR(default_rotation=0) # force unrotated (bypasses auto-orientation)
engine = OneOCR(max_lines=50) # limit output linesengine.recognize(image: PIL.Image, language=None, rotation=None, max_side=None, score_threshold=None, link_threshold=None) → OcrResult
Run OCR on a PIL Image object (any mode, any size 50–10 000 px). You can optionally override defaults on a per-call basis.
engine.recognize_file(path, language=None, rotation=None, max_side=None, score_threshold=None, link_threshold=None) → OcrResult
Convenience wrapper that opens the file and calls recognize().
| Field | Type | Description |
|---|---|---|
.lines |
list[Line] |
All recognized text lines |
.image_angle |
float |
Page rotation in degrees |
.full_text |
str (property) |
All lines joined by \n |
| Field | Type | Description |
|---|---|---|
.text |
str |
Full text of the line |
.bbox |
Quad |
Quadrilateral bounding box |
.style |
int |
0 = horizontal, 1 = vertical |
.words |
list[Word] |
Individual words |
| Field | Type | Description |
|---|---|---|
.text |
str |
Word text |
.bbox |
Quad |
Quadrilateral bounding box |
.confidence |
float |
Recognition confidence 0–1 |
Four corner points (x1,y1)…(x4,y4) clockwise from top-left.
quad.as_rect() # → (x, y, width, height) axis-aligned AABBoneocr.onemodel is an encrypted container. The DLL decrypts it internally using a hardcoded key before loading each sub-model via ONNX Runtime's CreateSessionFromArray C API. For a detailed step-by-step reverse engineering analysis of how this memory hooking and decryption works, see WRITEUP.md.
utils/extract_all.py intercepts this at runtime by:
- Loading
onnxruntime.dllinto the Python process. - Patching the API table of ONNX Runtime in-memory to hook
CreateSessionFromArrayandRun. - Loading
oneocr.dllto trigger model decryption. - Saving the decrypted model buffers and automatically reconstructing vocabulary files.
This technique requires no manual disassembly and works against future updates.
Note
Windows users: Docker is not needed at all. Run python utils/prepare_files.py natively — it handles everything automatically.
Docker is provided for Linux users who need to extract the models without a Windows machine.
Warning
This workflow uses Wine to emulate the Windows DLL decryption pipeline inside a Linux container. It is experimental and may break with future Wine or Snipping Tool updates.
docker build -f docker/Dockerfile.prepare -t oneocr-prepare .docker run --rm -v "$(pwd)/models:/output" oneocr-prepareAfter completion, the models/ directory will contain all ONNX sub-models and vocabulary files. The Docker image and cache can then be safely deleted — they are no longer needed.
# Free up disk space after extraction (~3-5 GB)
docker rmi oneocr-prepare
docker system pruneIf you prefer a containerized inference environment (e.g. for server deployment), you can build a standard Linux image after the models/ folder is populated:
docker build -f docker/Dockerfile -t oneocr-onnx-python .
docker run --rm -v "/path/to/images:/images" oneocr-onnx-python /images/screenshot.png- b1tg's Research Post on Win11 OneOCR — Detailed reverse engineering writeup of the OneOCR DLL engine.
- b1tg's win11-oneocr Repository — Original C++ proof-of-concept DLL hooking and extraction codebase.
- AuroraWright's oneocr Wrapper — Python-based OneOCR pipeline loader and model wrapper.
- Software License: The Python codebase, scripts, and documentation in this repository are licensed under the MIT License.
- Legal Disclaimer: This project is created for educational and research purposes only. This repository does not host or distribute any Microsoft proprietary binaries, DLLs, or neural network models. All proprietary assets must be acquired and decrypted locally by the end-user. All trademarks and technologies remain the property of their respective owners.