Skip to content

About

基于 Node.js 的截图内容识别系统

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Screenshot OCR

基于 Node.js 的截图内容识别系统,封装 Tesseract.js v7,支持中英文混合识别、图片预处理、区域识别和批量并发处理。同时提供程序化 API和 HTTP REST API 两种调用方式。

功能特性

  • 中英文混合识别 — 内置 chi_sim(简体中文)和 eng(英文)语言包,本地离线识别
  • 结构化输出 — Block → Paragraph → Line → Word → Symbol 五级层级,含文本、置信度、坐标
  • 区域识别 — 通过 Tesseract 原生 rectangle 参数零拷贝裁剪,只识别指定区域
  • 图片预处理 — 基于 sharp 的可配置管道:灰度化 → 直方图归一化 → 二值化 → 去噪 → 锐化
  • 批量并发处理 — Worker 池管理,可控并发数,单个失败不影响整体
  • 多输入格式 — 支持本地文件路径、Buffer、Base64 Data URI
  • Worker 池 — 惰性初始化,自动调度,支持超时控制和优雅关闭

快速开始

环境要求

  • Node.js >= 18
  • 无需安装 Tesseract 本体(tesseract.js 内置 WASM 引擎)

安装

npm install

启动 HTTP 服务

npm start
# 或
node src/index.js
╔══════════════════════════════════════════════════════╗
║   Screenshot OCR API Server                          ║
║   http://localhost:3000                              ║
║                                                      ║
║   POST /api/recognize  — 单图 OCR                    ║
║   POST /api/batch      — 批量 OCR                   ║
║   POST /api/region     — 区域 OCR                   ║
║   GET  /api/health     — 服务状态                   ║
╚══════════════════════════════════════════════════════╝

程序化调用

示例说明

本示例演示如何在 Node.js 应用中集成 Screenshot OCR 进行文本识别。涵盖初始化配置、图片识别、结果处理和资源释放的完整流程。

const ocr = require('./src');

// 配置(可选,在首次调用前设置)
// 建议在应用启动时配置一次,用于全局设置参数
ocr.configure({
  poolSize: 2,              // 并行 Worker 数,根据 CPU 核心数调整
  languages: 'chi_sim+eng', // 识别语言,支持 chi_sim、eng、osd 等
  jobTimeout: 120000,       // 单任务超时时间,毫秒为单位
});

// 识别单张截图
const result = await ocr.recognize('./screenshot.png', {
  preprocess: { grayscale: true, threshold: 128 },  // 预处理管道
  output: { blocks: true },                          // 返回结构化数据
});

// 处理识别结果
console.log(result.text);       // 纯文本:"识别到的完整文本"
console.log(result.confidence); // 整体置信度:0-100
console.log(result.blocks);     // 结构化层级:Block → Paragraph → Line → Word → Symbol
console.log(result.timing);     // 性能数据:{ preprocessMs, ocrMs, totalMs }

// 用完后关闭,释放 Worker 和系统资源
await ocr.shutdown();

程序化 API 参考

ocr.configure(options)

全局配置 OCR 运行时参数。应在首次调用 recognize / recognizeBatch / recognizeRegion 前执行,配置将影响所有后续操作。

参数 类型 默认值 说明
poolSize number 2 Worker 线程池大小,建议值为 CPU 核心数的 50-100%
languages string 'chi_sim+eng' Tesseract 语言代码,多语言用 + 连接(如 chi_sim+eng+fra)
langPath string 项目根目录 traineddata 文件目录路径
jobTimeout number 120000 单任务超时时间(毫秒),超过此时间视为失败

ocr.recognize(input, options?)

识别单张图片。

const result = await ocr.recognize(input, {
  preprocess: {
    grayscale: true,    // 灰度化
    normalize: true,    // 直方图归一化
    threshold: 128,     // 二值化阈值 (0-255)
    denoise: { radius: 2, type: 'median' },  // 中值滤波去噪
    sharpen: false,     // 锐化
  },
  output: {
    text: true,         // 纯文本输出
    blocks: true,       // 结构化层级输出
    hocr: false,        // HOCR 格式
    tsv: false,         // TSV 格式
  },
  tesseract: {
    rotateAuto: false,  // 自动旋转检测
    psm: 3,             // 页面分割模式
  },
  minConfidence: 60,    // 最低置信度过滤
  flattenTo: 'word',    // 扁平化到指定层级: word | line | paragraph | block
});

返回值:

{
  text: string;                  // 完整识别文本
  confidence: number;            // 整体置信度 (0-100)
  version: string;               // Tesseract 版本
  blocks: Block[];               // 结构化层级
  blockCount: number;
  paragraphCount: number;
  lineCount: number;
  wordCount: number;
  symbolCount: number;
  timing: {
    preprocessMs: number;        // 预处理耗时
    ocrMs: number;               // OCR 耗时
    totalMs: number;             // 总耗时
  };
}

每个 Block / Paragraph / Line / Word / Symbol 均包含:

{
  type: 'word';                  // block | paragraph | line | word | symbol
  text: 'Hello';                 // 该元素文本
  confidence: 95;                // 置信度
  bbox: { x: 10, y: 20, width: 60, height: 20 };  // 包围盒
  // 子元素...
}

ocr.recognizeRegion(input, region, options?)

识别图片指定区域。

const result = await ocr.recognizeRegion('./screenshot.png', {
  left: 100, top: 200, width: 400, height: 300
}, {
  preprocess: { threshold: 128 },
  output: { blocks: true },
});

返回值比 recognize() 多一个 region 字段,包含实际使用的区域坐标和是否被裁剪的标志。

ocr.recognizeBatch(images, options?, onProgress?)

批量识别,内部并发控制。

const result = await ocr.recognizeBatch([
  { id: 'img1', input: './a.png' },
  { id: 'img2', input: buffer2 },
  { id: 'img3', input: 'data:image/png;base64,...' },
], {
  preprocess: { grayscale: true },
}, (percent, completed, total) => {
  console.log(`进度: ${percent}% (${completed}/${total})`);
});

返回值:

{
  total: 3,
  succeeded: 2,
  failed: 1,
  results: [
    { id: 'img1', success: true, data: { ... } },
    { id: 'img2', success: true, data: { ... } },
    { id: 'img3', success: false, error: { code: 'OCR_TIMEOUT', message: '...' } },
  ],
  timing: { totalMs: 2500 },
}

ocr.preprocessImage(input, preprocessOptions?)

仅执行图片预处理,返回处理后的 PNG Buffer,不进行 OCR。

const pngBuffer = await ocr.preprocessImage('./screenshot.png', {
  grayscale: true,
  threshold: 128,
  denoise: { radius: 2 },
});

ocr.getStatus()

const status = await ocr.getStatus();
// { totalWorkers: 2, queueLength: 0, initialized: true }

ocr.shutdown()

关闭所有 Worker 并释放资源。

HTTP API 参考

POST /api/recognize

单图识别。支持三种传入方式:

方式 1:multipart 文件上传

curl -X POST http://localhost:3000/api/recognize \
  -F "image=@screenshot.png" \
  -F 'preprocess={"grayscale":true,"threshold":128}'

方式 2:JSON body — 文件路径

curl -X POST http://localhost:3000/api/recognize \
  -H "Content-Type: application/json" \
  -d '{"image_path": "./screenshot.png", "preprocess": {"threshold": 128}}'

方式 3:JSON body — Base64

curl -X POST http://localhost:3000/api/recognize \
  -H "Content-Type: application/json" \
  -d '{"image": "data:image/png;base64,iVBORw0KGgo...", "output": {"blocks": true}}'

请求参数:

字段 类型 默认 说明
image string — Base64 Data URI 或纯 base64
image_path string — 本地文件路径
preprocess object — 预处理配置
output object {"blocks":true} 输出格式 {text, blocks, hocr, tsv}
tesseract object — Tesseract 参数 {rectangle, rotateAuto, psm}
minConfidence number 0 置信度过滤阈值
flattenTo string — 扁平化层级

成功响应 (200):

{
  "success": true,
  "data": {
    "text": "识别文本...",
    "confidence": 91,
    "blocks": [ ... ],
    "timing": { "preprocessMs": 18, "ocrMs": 79, "totalMs": 97 }
  }
}

错误响应:

{
  "success": false,
  "error": {
    "code": "OCR_INIT_FAILED",
    "message": "Failed to initialize OCR worker...",
    "details": {}
  }
}

POST /api/region

与 /api/recognize 参数相同,额外需要 region 字段:

curl -X POST http://localhost:3000/api/region \
  -H "Content-Type: application/json" \
  -d '{
    "image_path": "./screenshot.png",
    "region": { "left": 100, "top": 200, "width": 400, "height": 300 },
    "output": { "blocks": true }
  }'

POST /api/batch

curl -X POST http://localhost:3000/api/batch \
  -H "Content-Type: application/json" \
  -d '{
    "images": [
      { "id": "img1", "image_path": "./a.png" },
      { "id": "img2", "image": "data:image/png;base64,..." }
    ],
    "preprocess": { "grayscale": true },
    "output": { "text": true }
  }'

或上传多个文件:

curl -X POST http://localhost:3000/api/batch \
  -F "images=@a.png" \
  -F "images=@b.png"

GET /api/health

curl http://localhost:3000/api/health
{
  "status": "ok",
  "workers": { "total": 2, "queueLength": 0 },
  "uptime": 3600.5,
  "config": { "defaultLanguages": "chi_sim+eng", "poolSize": 2, "jobTimeout": 120000 }
}

环境变量

变量 默认值 说明
PORT 3000 HTTP 服务端口
POOL_SIZE 2 Worker 线程数
DEFAULT_LANGUAGES chi_sim+eng 默认识别语言
MAX_FILE_SIZE 10485760 (10MB) 上传文件大小限制
JOB_TIMEOUT 120000 单任务超时 (ms)

错误码

错误码 HTTP 状态 说明
INVALID_INPUT 400 输入无效:缺少图片、不支持格式
IMAGE_TOO_LARGE 413 图片超过大小限制
OCR_INIT_FAILED 500 Worker 初始化失败(traineddata 缺失等)
OCR_TIMEOUT 504 识别超时
PREPROCESS_FAILED 422 图片预处理失败
WORKER_POOL_EXHAUSTED 503 Worker 池耗尽
INTERNAL_ERROR 500 内部未知错误

图片预处理说明

预处理管道基于 sharp,可独立控制每一步:

原始图片 → [灰度化] → [归一化] → [二值化] → [去噪] → [锐化] → PNG Buffer → Tesseract

推荐配置:

// 截图/UI 文字(推荐)
{ grayscale: true, normalize: true, threshold: 128, denoise: { radius: 2, type: 'median' } }

// 文档/扫描件
{ grayscale: true, normalize: true, threshold: 140, denoise: { radius: 1, type: 'median' }, sharpen: true }

// 不使用预处理
undefined  // 或 false

提供了两个快捷预设:

const { ImagePreprocessor } = require('./src');
ImagePreprocessor.screenshotPreset();  // 截图优化
ImagePreprocessor.documentPreset();    // 文档优化

项目结构

src/
├── index.js                     # 入口:模块导出 + HTTP 服务引导
├── core/
│   ├── OcrEngine.js             # Tesseract Worker 封装
│   ├── WorkerPool.js            # Worker 池 / 调度器
│   └── StructuredOutput.js      # 原始结果 → 结构化 JSON 转换
├── preprocessing/
│   └── ImagePreprocessor.js     # sharp 预处理管道
├── services/
│   ├── RecognitionService.js    # 单图 / 批量识别编排
│   └── RegionService.js         # 区域识别编排
├── server/
│   ├── app.js                   # Express 应用
│   ├── config.js                # 服务配置
│   ├── routes/
│   │   ├── recognize.js         # POST /api/recognize
│   │   ├── batch.js             # POST /api/batch
│   │   ├── region.js            # POST /api/region
│   │   └── health.js            # GET /api/health
│   └── middleware/
│       ├── upload.js            # multer 文件上传
│       ├── validate.js          # 请求校验
│       └── errorHandler.js      # 全局错误处理
└── utils/
    ├── OcrError.js              # 自定义错误类
    ├── paths.js                 # 路径工具
    └── imageLoader.js           # 图片加载(统一多输入格式)

依赖

包 版本 用途
tesseract.js ^7.0.0 OCR 引擎 (WASM)
sharp ^0.33.0 图片预处理
express ^4.21.0 HTTP 服务框架
multer ^1.4.5 文件上传处理
uuid ^10.0.0 Job ID 生成

License

Apache License 2.0

Copyright 2024-present Screenshot OCR Contributors

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

About

基于 Node.js 的截图内容识别系统

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages