52 KiB
PDFスキャン監視・自動振り分けツール Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: 既存のCLI型OCR氏名/日付抽出ツールを、バッチダブルクリック起動・固定フォルダ構成・同梱Python完結配布パッケージ型に全面改修する。
Architecture: box_detector.py(テンプレート枠検出)はそのまま流用。設定読込・ファイル操作・ロック管理・ログ出力を責任ごとに独立モジュールへ分離し、extract_and_rename.py はそれらを組み合わせるオーケストレーション層とする。配布用のpython/・tools/(tesseract/poppler)はビルドスクリプトが別途生成する。
Tech Stack: Python 3.11系(配布用embeddable、開発機は3.12.1で代用)、opencv-python、pdf2image、pytesseract、pillow、numpy、pytest(開発依存)、Windows batファイル。
Global Constraints
- 対象OS: Windows専用(
msvcrt・ctypes.windll使用前提) - 配布用Python: 3.11.9 embeddable amd64(バージョン固定)。開発・テストは開発機のシステムPython 3.12.1を使う(別物、配布物には含めない)
- pipパッケージ: opencv-python, pdf2image, pytesseract, pillow, numpy
- 外部バイナリ: tesseract-ocr(UB-Mannheim版 5.4.0.20240606固定)、poppler for Windows(poppler-windows 24.02.0-0固定)
- ログ出力エンコーディング: UTF-8 BOM付き(
utf-8-sig)、出力先ログ/YYYY-MM-DD.log - ファイル名衝突回避ルール:
stem(2).suffix,stem(3).suffix... の形式(既存の_1, _2形式ではない) config.txtは必須ファイル。存在しなければエラー終了(自動生成しない)- リポジトリ内のコード配置は
スクリプト/フォルダ配下(OCR仕分け実行.batとconfig.txtのみリポジトリ直下) - 全てのファイル・フォルダ・ログの文言は日本語
Task 1: リポジトリ構造整備
Files:
- Create:
スクリプト/(新規フォルダ) - Move:
box_detector.py→スクリプト/box_detector.py - Move:
extract_and_rename.py→スクリプト/extract_and_rename.py - Create:
.gitignore - Create:
スクリプト/tests/conftest.py - Create:
スクリプト/tests/test_smoke.py
Interfaces:
-
Produces:
スクリプト/tests/conftest.pyがsys.pathにスクリプト/を追加する。以降の全テストファイルはこれを前提にfrom config_loader import ...のような直接importができる。 -
Step 1: 既存ファイルをスクリプト/フォルダへ移動
cd "c:/Users/k.nogi/#GitHub/ken_nogi/PythonProj/ScanOCR"
mkdir -p スクリプト
git mv box_detector.py スクリプト/box_detector.py
git mv extract_and_rename.py スクリプト/extract_and_rename.py
- Step 2: .gitignore作成
.gitignore に以下を書く:
dist/
__pycache__/
*.pyc
.pytest_cache/
.lock
ログ/
- Step 3: pytestインストール・conftest.py作成
python -m pip install pytest
mkdir -p スクリプト/tests
スクリプト/tests/conftest.py:
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
スクリプト/tests/test_smoke.py:
def test_smoke():
assert True
- Step 4: テスト実行確認
Run: python -m pytest スクリプト/tests -v
Expected: test_smoke.py::test_smoke PASSED、1 passed
- Step 5: コミット
git add .gitignore スクリプト/
git commit -m "chore: スクリプト配置をスクリプト/フォルダへ整理しテスト基盤を追加"
Task 2: config_loader.py(config.txt読込)
Files:
- Create:
スクリプト/config_loader.py - Test:
スクリプト/tests/test_config_loader.py
Interfaces:
-
Produces:
class ConfigError(Exception)REQUIRED_KEYS: tuple[str, ...]load_config(config_path: Path) -> dict[str, str]— 全項目を文字列のまま返す。数値変換は呼び出し側の責務
-
Step 1: 失敗するテストを書く
スクリプト/tests/test_config_loader.py:
import pytest
from pathlib import Path
from config_loader import ConfigError, load_config
VALID_CONTENT = """\
スクリプトフォルダ=スクリプト
テンプレートフォルダ=テンプレート
スキャンフォルダ=スキャン
アウトプットフォルダ=アウトプット
成功フォルダ=成功
失敗フォルダ=失敗
ログフォルダ=ログ
margin=0.10
DPI=300
ファイル安定待ち秒=1
ファイル安定待ちリトライ回数=5
"""
def test_load_config_missing_file_raises(tmp_path):
missing_path = tmp_path / "config.txt"
with pytest.raises(ConfigError):
load_config(missing_path)
def test_load_config_parses_all_keys(tmp_path):
config_path = tmp_path / "config.txt"
config_path.write_text(VALID_CONTENT, encoding="utf-8")
config = load_config(config_path)
assert config["スクリプトフォルダ"] == "スクリプト"
assert config["テンプレートフォルダ"] == "テンプレート"
assert config["margin"] == "0.10"
assert config["ファイル安定待ちリトライ回数"] == "5"
def test_load_config_missing_required_key_raises(tmp_path):
config_path = tmp_path / "config.txt"
config_path.write_text("スクリプトフォルダ=スクリプト\n", encoding="utf-8")
with pytest.raises(ConfigError):
load_config(config_path)
def test_load_config_ignores_blank_lines_and_comments(tmp_path):
config_path = tmp_path / "config.txt"
content = "# comment\n\n" + VALID_CONTENT
config_path.write_text(content, encoding="utf-8")
config = load_config(config_path)
assert config["スクリプトフォルダ"] == "スクリプト"
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_config_loader.py -v
Expected: ModuleNotFoundError: No module named 'config_loader' で全件FAIL
- Step 3: config_loader.py実装
スクリプト/config_loader.py:
# -*- coding: utf-8 -*-
"""config.txt(key=value形式)読込モジュール。"""
from __future__ import annotations
from pathlib import Path
class ConfigError(Exception):
pass
REQUIRED_KEYS = (
"スクリプトフォルダ",
"テンプレートフォルダ",
"スキャンフォルダ",
"アウトプットフォルダ",
"成功フォルダ",
"失敗フォルダ",
"ログフォルダ",
"margin",
"DPI",
"ファイル安定待ち秒",
"ファイル安定待ちリトライ回数",
)
def load_config(config_path: Path) -> dict[str, str]:
if not config_path.exists():
raise ConfigError(f"config.txt が見つかりません: {config_path}")
config: dict[str, str] = {}
for raw_line in config_path.read_text(encoding="utf-8").splitlines():
line = raw_line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
config[key.strip()] = value.strip()
missing = [key for key in REQUIRED_KEYS if key not in config]
if missing:
raise ConfigError(
f"config.txt に必須項目が不足しています: {', '.join(missing)}"
)
return config
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_config_loader.py -v
Expected: 4 passed
- Step 5: コミット
git add スクリプト/config_loader.py スクリプト/tests/test_config_loader.py
git commit -m "feat: config.txt読込モジュールを追加"
Task 3: file_ops.py(ファイル名衝突回避コピー・移動)
Files:
- Create:
スクリプト/file_ops.py - Test:
スクリプト/tests/test_file_ops.py
Interfaces:
-
Produces:
resolve_unique_path(dest_dir: Path, stem: str, suffix: str) -> Pathcopy_with_unique_name(src: Path, dest_dir: Path, stem: str, suffix: str) -> Pathmove_with_unique_name(src: Path, dest_dir: Path, stem: str, suffix: str) -> Path
-
Step 1: 失敗するテストを書く
スクリプト/tests/test_file_ops.py:
from pathlib import Path
from file_ops import copy_with_unique_name, move_with_unique_name, resolve_unique_path
def test_resolve_unique_path_no_collision(tmp_path):
result = resolve_unique_path(tmp_path, "氏名A_20260101", ".pdf")
assert result == tmp_path / "氏名A_20260101.pdf"
def test_resolve_unique_path_with_collision(tmp_path):
(tmp_path / "氏名A_20260101.pdf").write_bytes(b"x")
result = resolve_unique_path(tmp_path, "氏名A_20260101", ".pdf")
assert result == tmp_path / "氏名A_20260101(2).pdf"
def test_resolve_unique_path_with_multiple_collisions(tmp_path):
(tmp_path / "氏名A_20260101.pdf").write_bytes(b"x")
(tmp_path / "氏名A_20260101(2).pdf").write_bytes(b"x")
result = resolve_unique_path(tmp_path, "氏名A_20260101", ".pdf")
assert result == tmp_path / "氏名A_20260101(3).pdf"
def test_copy_with_unique_name_creates_dest_dir(tmp_path):
src = tmp_path / "src.pdf"
src.write_bytes(b"content")
dest_dir = tmp_path / "output" / "氏名A"
result = copy_with_unique_name(src, dest_dir, "氏名A_20260101", ".pdf")
assert result == dest_dir / "氏名A_20260101.pdf"
assert result.read_bytes() == b"content"
assert src.exists() # コピーなので元ファイルは残る
def test_move_with_unique_name_moves_source(tmp_path):
src = tmp_path / "src.pdf"
src.write_bytes(b"content")
dest_dir = tmp_path / "success"
result = move_with_unique_name(src, dest_dir, "src", ".pdf")
assert result == dest_dir / "src.pdf"
assert result.read_bytes() == b"content"
assert not src.exists() # 移動なので元ファイルは消える
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_file_ops.py -v
Expected: ModuleNotFoundError: No module named 'file_ops' で全件FAIL
- Step 3: file_ops.py実装(衝突回避部分)
スクリプト/file_ops.py:
# -*- coding: utf-8 -*-
"""ファイル名衝突回避コピー・移動、ファイルサイズ安定待ちモジュール。"""
from __future__ import annotations
import shutil
import time
from pathlib import Path
from typing import Callable
def resolve_unique_path(dest_dir: Path, stem: str, suffix: str) -> Path:
candidate = dest_dir / f"{stem}{suffix}"
if not candidate.exists():
return candidate
counter = 2
while True:
candidate = dest_dir / f"{stem}({counter}){suffix}"
if not candidate.exists():
return candidate
counter += 1
def copy_with_unique_name(src: Path, dest_dir: Path, stem: str, suffix: str) -> Path:
dest_dir.mkdir(parents=True, exist_ok=True)
dest_path = resolve_unique_path(dest_dir, stem, suffix)
shutil.copy2(str(src), str(dest_path))
return dest_path
def move_with_unique_name(src: Path, dest_dir: Path, stem: str, suffix: str) -> Path:
dest_dir.mkdir(parents=True, exist_ok=True)
dest_path = resolve_unique_path(dest_dir, stem, suffix)
shutil.move(str(src), str(dest_path))
return dest_path
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_file_ops.py -v
Expected: 5 passed
- Step 5: コミット
git add スクリプト/file_ops.py スクリプト/tests/test_file_ops.py
git commit -m "feat: ファイル名衝突回避コピー・移動モジュールを追加"
Task 4: file_ops.py(ファイルサイズ安定待ち)
Files:
- Modify:
スクリプト/file_ops.py(Task 3で作成したファイルに追記) - Test:
スクリプト/tests/test_file_ops.py(追記)
Interfaces:
-
Consumes: なし(Task 3の関数とは独立)
-
Produces:
wait_until_stable(path: Path, interval_sec: float, retries: int, sleep_func: Callable[[float], None] = time.sleep) -> bool- ファイルが存在しなければ即
False interval_sec秒待って直前サイズと比較、一致すればTrueretries回試行しても一致しなければFalse
- ファイルが存在しなければ即
-
Step 1: 失敗するテストを書く
スクリプト/tests/test_file_ops.py に追記:
from file_ops import wait_until_stable
def test_wait_until_stable_missing_file_returns_false(tmp_path):
missing = tmp_path / "missing.pdf"
result = wait_until_stable(missing, interval_sec=0, retries=3, sleep_func=lambda s: None)
assert result is False
def test_wait_until_stable_stable_file_returns_true(tmp_path):
path = tmp_path / "stable.pdf"
path.write_bytes(b"1234")
result = wait_until_stable(path, interval_sec=0, retries=3, sleep_func=lambda s: None)
assert result is True
def test_wait_until_stable_growing_file_returns_false(tmp_path):
path = tmp_path / "growing.pdf"
path.write_bytes(b"1")
call_count = {"n": 0}
def fake_sleep(_seconds):
call_count["n"] += 1
path.write_bytes(b"1" * (call_count["n"] + 1))
result = wait_until_stable(path, interval_sec=0, retries=3, sleep_func=fake_sleep)
assert result is False
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_file_ops.py -v -k wait_until_stable
Expected: ImportError: cannot import name 'wait_until_stable' で全件FAIL
- Step 3: wait_until_stable実装
スクリプト/file_ops.py の末尾に追記:
def wait_until_stable(
path: Path,
interval_sec: float,
retries: int,
sleep_func: Callable[[float], None] = time.sleep,
) -> bool:
if not path.exists():
return False
previous_size = path.stat().st_size
for _ in range(retries):
sleep_func(interval_sec)
if not path.exists():
return False
current_size = path.stat().st_size
if current_size == previous_size:
return True
previous_size = current_size
return False
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_file_ops.py -v
Expected: 8 passed
- Step 5: コミット
git add スクリプト/file_ops.py スクリプト/tests/test_file_ops.py
git commit -m "feat: ファイルサイズ安定待ちロジックを追加"
Task 5: lock_manager.py(二重起動防止ロック)
Files:
- Create:
スクリプト/lock_manager.py - Test:
スクリプト/tests/test_lock_manager.py
Interfaces:
-
Produces:
class LockAcquisitionError(Exception)acquire_lock(lock_path: Path, pid_checker: Callable[[int], bool] = _is_pid_running_windows) -> Nonerelease_lock(lock_path: Path) -> None
-
Step 1: 失敗するテストを書く
スクリプト/tests/test_lock_manager.py:
import os
import pytest
from lock_manager import LockAcquisitionError, acquire_lock, release_lock
def test_acquire_lock_creates_file_with_pid(tmp_path):
lock_path = tmp_path / ".lock"
acquire_lock(lock_path, pid_checker=lambda pid: True)
assert lock_path.exists()
assert lock_path.read_text(encoding="utf-8").strip() == str(os.getpid())
def test_acquire_lock_raises_when_existing_pid_running(tmp_path):
lock_path = tmp_path / ".lock"
lock_path.write_text("12345", encoding="utf-8")
with pytest.raises(LockAcquisitionError):
acquire_lock(lock_path, pid_checker=lambda pid: True)
def test_acquire_lock_replaces_stale_lock(tmp_path):
lock_path = tmp_path / ".lock"
lock_path.write_text("12345", encoding="utf-8")
acquire_lock(lock_path, pid_checker=lambda pid: False)
assert lock_path.read_text(encoding="utf-8").strip() == str(os.getpid())
def test_release_lock_removes_file(tmp_path):
lock_path = tmp_path / ".lock"
lock_path.write_text("12345", encoding="utf-8")
release_lock(lock_path)
assert not lock_path.exists()
def test_release_lock_missing_file_no_error(tmp_path):
lock_path = tmp_path / ".lock"
release_lock(lock_path) # 例外が出ないことを確認するだけ
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_lock_manager.py -v
Expected: ModuleNotFoundError: No module named 'lock_manager' で全件FAIL
- Step 3: lock_manager.py実装
スクリプト/lock_manager.py:
# -*- coding: utf-8 -*-
"""二重起動防止ロックファイル管理モジュール(Windows専用)。"""
from __future__ import annotations
import ctypes
import os
from pathlib import Path
from typing import Callable
class LockAcquisitionError(Exception):
pass
def _is_pid_running_windows(pid: int) -> bool:
PROCESS_QUERY_LIMITED_INFORMATION = 0x1000
handle = ctypes.windll.kernel32.OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION, False, pid)
if handle:
ctypes.windll.kernel32.CloseHandle(handle)
return True
return False
def acquire_lock(
lock_path: Path,
pid_checker: Callable[[int], bool] = _is_pid_running_windows,
) -> None:
if lock_path.exists():
try:
recorded_pid = int(lock_path.read_text(encoding="utf-8").strip())
except ValueError:
recorded_pid = None
if recorded_pid is not None and pid_checker(recorded_pid):
raise LockAcquisitionError(f"既に実行中です(PID={recorded_pid})")
lock_path.unlink()
lock_path.write_text(str(os.getpid()), encoding="utf-8")
def release_lock(lock_path: Path) -> None:
if lock_path.exists():
lock_path.unlink()
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_lock_manager.py -v
Expected: 5 passed
- Step 5: コミット
git add スクリプト/lock_manager.py スクリプト/tests/test_lock_manager.py
git commit -m "feat: 二重起動防止ロック管理モジュールを追加"
Task 6: logger.py(実行ログ出力)
Files:
- Create:
スクリプト/logger.py - Test:
スクリプト/tests/test_logger.py
Interfaces:
-
Produces:
LOG_ENCODING = "utf-8-sig"log_path_for_today(log_dir: Path, now_func: Callable[[], datetime] = datetime.now) -> Pathappend_log(log_dir: Path, result: str, original_filename: str, name: str | None, date: str | None, note: str = "", now_func: Callable[[], datetime] = datetime.now) -> None
-
Step 1: 失敗するテストを書く
スクリプト/tests/test_logger.py:
from datetime import datetime
from logger import LOG_ENCODING, append_log, log_path_for_today
def fixed_now():
return datetime(2026, 8, 2, 10, 30, 15)
def test_log_path_for_today_uses_now_func(tmp_path):
result = log_path_for_today(tmp_path, now_func=fixed_now)
assert result == tmp_path / "2026-08-02.log"
def test_append_log_creates_dir_and_file(tmp_path):
log_dir = tmp_path / "ログ"
append_log(
log_dir,
result="成功",
original_filename="scan001.pdf",
name="山田太郎",
date="20260802",
note="",
now_func=fixed_now,
)
log_path = log_dir / "2026-08-02.log"
assert log_path.exists()
content = log_path.read_text(encoding=LOG_ENCODING)
assert "[2026-08-02 10:30:15]" in content
assert "結果=成功" in content
assert "元ファイル=scan001.pdf" in content
assert "氏名=山田太郎" in content
assert "日付=20260802" in content
def test_append_log_handles_none_name_and_date(tmp_path):
log_dir = tmp_path / "ログ"
append_log(
log_dir,
result="失敗",
original_filename="scan002.pdf",
name=None,
date=None,
note="抽出不十分",
now_func=fixed_now,
)
content = (log_dir / "2026-08-02.log").read_text(encoding=LOG_ENCODING)
assert "氏名=" in content
assert "日付=" in content
assert "備考=抽出不十分" in content
def test_append_log_appends_multiple_lines(tmp_path):
log_dir = tmp_path / "ログ"
append_log(log_dir, "成功", "a.pdf", "氏名A", "20260801", "", now_func=fixed_now)
append_log(log_dir, "成功", "b.pdf", "氏名B", "20260802", "", now_func=fixed_now)
content = (log_dir / "2026-08-02.log").read_text(encoding=LOG_ENCODING)
assert content.count("結果=成功") == 2
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_logger.py -v
Expected: ModuleNotFoundError: No module named 'logger' で全件FAIL
- Step 3: logger.py実装
スクリプト/logger.py:
# -*- coding: utf-8 -*-
"""実行ログ出力モジュール。日付別ファイル・UTF-8 BOM付き。"""
from __future__ import annotations
from datetime import datetime
from pathlib import Path
from typing import Callable
LOG_ENCODING = "utf-8-sig"
def log_path_for_today(
log_dir: Path, now_func: Callable[[], datetime] = datetime.now
) -> Path:
return log_dir / f"{now_func().strftime('%Y-%m-%d')}.log"
def append_log(
log_dir: Path,
result: str,
original_filename: str,
name: str | None,
date: str | None,
note: str = "",
now_func: Callable[[], datetime] = datetime.now,
) -> None:
log_dir.mkdir(parents=True, exist_ok=True)
now = now_func()
line = (
f"[{now.strftime('%Y-%m-%d %H:%M:%S')}] "
f"結果={result} | 元ファイル={original_filename} | "
f"氏名={name or ''} | 日付={date or ''} | 備考={note}\n"
)
path = log_path_for_today(log_dir, now_func)
with open(path, "a", encoding=LOG_ENCODING) as f:
f.write(line)
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_logger.py -v
Expected: 4 passed
- Step 5: コミット
git add スクリプト/logger.py スクリプト/tests/test_logger.py
git commit -m "feat: 実行ログ出力モジュールを追加"
Task 7: extract_and_rename.py — 初期化系ヘルパー関数(テンプレート選定・依存チェック)
Files:
- Modify:
スクリプト/extract_and_rename.py - Test:
スクリプト/tests/test_extract_and_rename.py
Interfaces:
- Consumes: なし
- Produces:
TEMPLATE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".bmp"}select_template_file(template_dir: Path) -> PathREQUIRED_MODULES = ("cv2", "numpy", "PIL", "pytesseract", "pdf2image")check_dependencies(modules: tuple[str, ...] = REQUIRED_MODULES) -> list[str]check_external_tools(tesseract_exe: Path, poppler_exe: Path) -> list[str]
この時点では既存の FieldBox / load_field_boxes / ocr_region / clean_name / clean_date / process_pdf はそのまま残す(Task 8で main() を書き換える際に整理する)。既存の safe_move 関数と旧 main()(argparse版)はこのTaskではまだ削除しない。
- Step 1: 失敗するテストを書く
スクリプト/tests/test_extract_and_rename.py:
from pathlib import Path
from extract_and_rename import (
check_dependencies,
check_external_tools,
select_template_file,
)
def test_select_template_file_picks_first_by_name(tmp_path):
(tmp_path / "b_template.png").write_bytes(b"x")
(tmp_path / "a_template.png").write_bytes(b"x")
(tmp_path / "note.txt").write_bytes(b"x") # 対象外拡張子
result = select_template_file(tmp_path)
assert result == tmp_path / "a_template.png"
def test_select_template_file_no_candidates_raises(tmp_path):
(tmp_path / "note.txt").write_bytes(b"x")
try:
select_template_file(tmp_path)
assert False, "FileNotFoundErrorが送出されるべき"
except FileNotFoundError:
pass
def test_check_dependencies_detects_missing_module():
result = check_dependencies(("this_module_does_not_exist_xyz",))
assert result == ["this_module_does_not_exist_xyz"]
def test_check_dependencies_detects_installed_module():
result = check_dependencies(("os",))
assert result == []
def test_check_external_tools_detects_missing(tmp_path):
tesseract = tmp_path / "tools" / "tesseract" / "tesseract.exe"
poppler = tmp_path / "tools" / "poppler" / "pdftoppm.exe"
result = check_external_tools(tesseract, poppler)
assert str(tesseract) in result
assert str(poppler) in result
def test_check_external_tools_all_present(tmp_path):
tesseract = tmp_path / "tesseract.exe"
poppler = tmp_path / "pdftoppm.exe"
tesseract.write_bytes(b"x")
poppler.write_bytes(b"x")
result = check_external_tools(tesseract, poppler)
assert result == []
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_extract_and_rename.py -v
Expected: ImportError(select_template_file 等が存在しない)で全件FAIL
- Step 3: ヘルパー関数を実装
スクリプト/extract_and_rename.py の先頭 import 群の直後に追記(既存コードは変更しない):
import importlib
TEMPLATE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".bmp"}
def select_template_file(template_dir: Path) -> Path:
candidates = sorted(
p for p in template_dir.iterdir()
if p.is_file() and p.suffix.lower() in TEMPLATE_EXTENSIONS
)
if not candidates:
raise FileNotFoundError(
f"テンプレートフォルダに画像ファイルが見つかりません: {template_dir}"
)
return candidates[0]
REQUIRED_MODULES = ("cv2", "numpy", "PIL", "pytesseract", "pdf2image")
def check_dependencies(modules: tuple[str, ...] = REQUIRED_MODULES) -> list[str]:
missing = []
for mod in modules:
try:
importlib.import_module(mod)
except ImportError:
missing.append(mod)
return missing
def check_external_tools(tesseract_exe: Path, poppler_exe: Path) -> list[str]:
missing = []
if not tesseract_exe.exists():
missing.append(str(tesseract_exe))
if not poppler_exe.exists():
missing.append(str(poppler_exe))
return missing
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_extract_and_rename.py -v
Expected: 5 passed
- Step 5: コミット
git add スクリプト/extract_and_rename.py スクリプト/tests/test_extract_and_rename.py
git commit -m "feat: テンプレート選定・依存チェックヘルパーを追加"
Task 8: extract_and_rename.py — キュー処理ループ(run_queue)
Files:
- Modify:
スクリプト/extract_and_rename.py - Test:
スクリプト/tests/test_extract_and_rename.py(追記)
Interfaces:
-
Consumes:
file_ops.copy_with_unique_name,file_ops.move_with_unique_name,file_ops.wait_until_stablelogger.append_log- 既存
process_pdf(pdf_path: Path, boxes: list[FieldBox], margin: float) -> tuple[str | None, str | None]
-
Produces:
key_pressed() -> bool(msvcrt.kbhit()の薄いラッパー)run_queue(scan_dir: Path, output_dir: Path, success_dir: Path, failed_dir: Path, log_dir: Path, boxes: list[FieldBox], margin: float, stable_wait_sec: float, stable_retries: int, process_pdf_func=process_pdf, key_check=key_pressed, sleep_func=time.sleep) -> None
-
Step 1: 失敗するテストを書く
スクリプト/tests/test_extract_and_rename.py に追記:
from extract_and_rename import run_queue
def _make_dirs(tmp_path):
dirs = {
"scan": tmp_path / "スキャン",
"output": tmp_path / "アウトプット",
"success": tmp_path / "成功",
"failed": tmp_path / "失敗",
"log": tmp_path / "ログ",
}
dirs["scan"].mkdir()
return dirs
def test_run_queue_success_moves_and_copies(tmp_path):
dirs = _make_dirs(tmp_path)
pdf = dirs["scan"] / "scan001.pdf"
pdf.write_bytes(b"dummy")
def fake_process_pdf(path, boxes, margin):
return ("氏名A", "20260802")
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=fake_process_pdf,
key_check=lambda: False,
sleep_func=lambda s: None,
)
assert (dirs["output"] / "氏名A" / "氏名A_20260802.pdf").exists()
assert (dirs["success"] / "scan001.pdf").exists()
assert not pdf.exists()
log_content = (dirs["log"] / f"{__import__('datetime').datetime.now().strftime('%Y-%m-%d')}.log").read_text(encoding="utf-8-sig")
assert "結果=成功" in log_content
def test_run_queue_missing_name_moves_to_failed(tmp_path):
dirs = _make_dirs(tmp_path)
pdf = dirs["scan"] / "scan002.pdf"
pdf.write_bytes(b"dummy")
def fake_process_pdf(path, boxes, margin):
return (None, "20260802")
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=fake_process_pdf,
key_check=lambda: False,
sleep_func=lambda s: None,
)
assert (dirs["failed"] / "scan002.pdf").exists()
assert not pdf.exists()
def test_run_queue_exception_moves_to_failed(tmp_path):
dirs = _make_dirs(tmp_path)
pdf = dirs["scan"] / "scan003.pdf"
pdf.write_bytes(b"dummy")
def fake_process_pdf(path, boxes, margin):
raise RuntimeError("OCRエラー")
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=fake_process_pdf,
key_check=lambda: False,
sleep_func=lambda s: None,
)
assert (dirs["failed"] / "scan003.pdf").exists()
def test_run_queue_stops_on_key_press(tmp_path):
dirs = _make_dirs(tmp_path)
pdf = dirs["scan"] / "scan004.pdf"
pdf.write_bytes(b"dummy")
calls = []
def fake_process_pdf(path, boxes, margin):
calls.append(path)
return ("氏名A", "20260802")
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=fake_process_pdf,
key_check=lambda: True, # 常に押下扱い
sleep_func=lambda s: None,
)
assert calls == []
assert pdf.exists() # 処理されず残る
def test_run_queue_skips_file_deleted_before_processing(tmp_path):
dirs = _make_dirs(tmp_path)
pdf = dirs["scan"] / "scan005.pdf"
pdf.write_bytes(b"dummy")
calls = []
def fake_process_pdf(path, boxes, margin):
calls.append(path)
return ("氏名A", "20260802")
def key_check_and_delete():
if pdf.exists():
pdf.unlink()
return False
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=fake_process_pdf,
key_check=key_check_and_delete,
sleep_func=lambda s: None,
)
assert calls == []
def test_run_queue_empty_scan_dir_returns_immediately(tmp_path):
dirs = _make_dirs(tmp_path)
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=lambda p, b, m: ("氏名A", "20260802"),
key_check=lambda: False,
sleep_func=lambda s: None,
)
# 例外が出ずに戻ることのみ確認
def test_run_queue_name_collision_gets_suffix(tmp_path):
dirs = _make_dirs(tmp_path)
existing_output_dir = dirs["output"] / "氏名A"
existing_output_dir.mkdir(parents=True)
(existing_output_dir / "氏名A_20260802.pdf").write_bytes(b"old")
pdf = dirs["scan"] / "scan006.pdf"
pdf.write_bytes(b"dummy")
run_queue(
dirs["scan"], dirs["output"], dirs["success"], dirs["failed"], dirs["log"],
boxes=[], margin=0.1, stable_wait_sec=0, stable_retries=1,
process_pdf_func=lambda p, b, m: ("氏名A", "20260802"),
key_check=lambda: False,
sleep_func=lambda s: None,
)
assert (existing_output_dir / "氏名A_20260802(2).pdf").exists()
- Step 2: テスト実行し失敗を確認
Run: python -m pytest スクリプト/tests/test_extract_and_rename.py -v -k run_queue
Expected: ImportError: cannot import name 'run_queue' で全件FAIL
- Step 3: run_queue実装
スクリプト/extract_and_rename.py の Task 7 で追加した関数群の直後に追記。合わせてファイル先頭の import 群に import time、import traceback、from file_ops import copy_with_unique_name, move_with_unique_name, wait_until_stable、from logger import append_log を追加する:
import time
import traceback
from file_ops import copy_with_unique_name, move_with_unique_name, wait_until_stable
from logger import append_log
def key_pressed() -> bool:
import msvcrt
return msvcrt.kbhit()
def run_queue(
scan_dir: Path,
output_dir: Path,
success_dir: Path,
failed_dir: Path,
log_dir: Path,
boxes: list[FieldBox],
margin: float,
stable_wait_sec: float,
stable_retries: int,
process_pdf_func=process_pdf,
key_check=key_pressed,
sleep_func=time.sleep,
) -> None:
while True:
queue = sorted(scan_dir.glob("*.pdf"))
if not queue:
break
pdf_path = queue[0]
if key_check():
append_log(
log_dir, "エラー", pdf_path.name, None, None,
"ユーザー操作により停止しました",
)
break
if not pdf_path.exists():
continue
if not wait_until_stable(pdf_path, stable_wait_sec, stable_retries, sleep_func):
append_log(
log_dir, "スキップ", pdf_path.name, None, None,
"サイズ不安定のためスキップ",
)
continue
try:
name, date = process_pdf_func(pdf_path, boxes, margin)
except Exception as e:
append_log(
log_dir, "エラー", pdf_path.name, None, None,
f"{type(e).__name__}: {e}\n{traceback.format_exc()}",
)
move_with_unique_name(pdf_path, failed_dir, pdf_path.stem, pdf_path.suffix)
continue
if not name or not date:
append_log(log_dir, "失敗", pdf_path.name, name, date, "")
move_with_unique_name(pdf_path, failed_dir, pdf_path.stem, pdf_path.suffix)
continue
copy_with_unique_name(pdf_path, output_dir / name, f"{name}_{date}", ".pdf")
move_with_unique_name(pdf_path, success_dir, pdf_path.stem, pdf_path.suffix)
append_log(log_dir, "成功", pdf_path.name, name, date, "")
- Step 4: テスト実行し成功を確認
Run: python -m pytest スクリプト/tests/test_extract_and_rename.py -v
Expected: 12 passed(Task 7の5件 + 本Taskの7件)
- Step 5: コミット
git add スクリプト/extract_and_rename.py スクリプト/tests/test_extract_and_rename.py
git commit -m "feat: キュー処理メインループ(run_queue)を追加"
Task 9: extract_and_rename.py — main()統合・旧CLI削除
Files:
- Modify:
スクリプト/extract_and_rename.py
Interfaces:
- Consumes:
config_loader.load_config,lock_manager.acquire_lock/release_lock/LockAcquisitionError,box_detector.detect_template_fields, Task 7/8で作った全ヘルパー - Produces:
main() -> None(配布パッケージ実行時のエントリーポイント)
このTaskはオーケストレーション層のみでユニットテストは書かない(下位関数は既にTask 2〜8でテスト済みのため)。既存の argparse ベースの旧 main() と safe_move 関数を削除し、固定フォルダ構成版に置き換える。
- Step 1: 旧mainとsafe_moveを削除
スクリプト/extract_and_rename.py から def safe_move(...) 関数全体と、旧 def main():(argparseを使う版)を削除する。if __name__ == "__main__": ブロックも一旦削除する(Step 2で再作成)。
- Step 2: 新main()を実装
ファイル末尾に追記:
from config_loader import load_config
from lock_manager import LockAcquisitionError, acquire_lock, release_lock
HOME_DIR = Path(__file__).resolve().parent.parent
RUNTIME_FOLDER_KEYS = (
"テンプレートフォルダ",
"スキャンフォルダ",
"アウトプットフォルダ",
"成功フォルダ",
"失敗フォルダ",
"ログフォルダ",
)
def main() -> None:
config_path = HOME_DIR / "config.txt"
try:
config = load_config(config_path)
except Exception as e:
print(f"設定エラー: {e}")
input("何かキーを押すと終了します...")
return
folders = {key: HOME_DIR / config[key] for key in RUNTIME_FOLDER_KEYS}
for folder in folders.values():
folder.mkdir(parents=True, exist_ok=True)
missing_modules = check_dependencies()
if missing_modules:
print(f"配布パッケージが壊れています。不足モジュール: {', '.join(missing_modules)}")
input("何かキーを押すと終了します...")
return
script_dir = HOME_DIR / config["スクリプトフォルダ"]
tesseract_exe = script_dir / "tools" / "tesseract" / "tesseract.exe"
poppler_exe = script_dir / "tools" / "poppler" / "pdftoppm.exe"
missing_tools = check_external_tools(tesseract_exe, poppler_exe)
if missing_tools:
print(f"配布パッケージが壊れています。不足ファイル: {', '.join(missing_tools)}")
input("何かキーを押すと終了します...")
return
lock_path = HOME_DIR / ".lock"
try:
acquire_lock(lock_path)
except LockAcquisitionError as e:
print(str(e))
input("何かキーを押すと終了します...")
return
try:
try:
template_path = select_template_file(folders["テンプレートフォルダ"])
except FileNotFoundError as e:
print(str(e))
return
print(f"テンプレート採用: {template_path.name}")
boxes = load_field_boxes(str(template_path))
print(f"検出フィールド: {[b.label for b in boxes]}")
margin = float(config["margin"])
stable_wait_sec = float(config["ファイル安定待ち秒"])
stable_retries = int(config["ファイル安定待ちリトライ回数"])
run_queue(
folders["スキャンフォルダ"],
folders["アウトプットフォルダ"],
folders["成功フォルダ"],
folders["失敗フォルダ"],
folders["ログフォルダ"],
boxes,
margin,
stable_wait_sec,
stable_retries,
)
finally:
release_lock(lock_path)
print("処理完了。")
input("何かキーを押すと終了します...")
if __name__ == "__main__":
main()
- Step 3: DPI定数の重複確認
ファイル冒頭の既存 DPI = 300 定数はそのまま残す(config.txt の DPI 項目は将来的な外部化用として保持するが、このTaskでは process_pdf 内の convert_from_path 呼び出しは既存の DPI 定数を使い続ける。config由来のDPI値を実際に反映するのは本Taskのスコープ外とし、次Task 10のREADMEに既知の制約として明記する)。
- Step 4: 全テスト実行し既存テストが壊れていないことを確認
Run: python -m pytest スクリプト/tests -v
Expected: 全件PASSED(Task 1〜8で書いた全テストが引き続き通る)
- Step 5: コミット
git add スクリプト/extract_and_rename.py
git commit -m "feat: 固定フォルダ構成・ロック・ログ統合のmain()に置き換え"
Task 10: OCR仕分け実行.bat作成
Files:
- Create:
OCR仕分け実行.bat
Interfaces:
-
Consumes:
config.txt(スクリプトフォルダ=行)、<スクリプトフォルダ>/python/python.exe、<スクリプトフォルダ>/extract_and_rename.py -
Step 1: バッチファイル作成
OCR仕分け実行.bat:
@echo off
chcp 65001 >nul
cd /d "%~dp0"
if not exist "config.txt" (
echo config.txt が見つかりません。配布パッケージが不完全です。
pause
exit /b 1
)
set "SCRIPT_DIR="
for /f "usebackq tokens=1,2 delims==" %%A in ("config.txt") do (
if "%%A"=="スクリプトフォルダ" set "SCRIPT_DIR=%%B"
)
if "%SCRIPT_DIR%"=="" (
echo config.txt に スクリプトフォルダ の設定が見つかりません。
pause
exit /b 1
)
set "PYTHON_EXE=%~dp0%SCRIPT_DIR%\python\python.exe"
set "MAIN_PY=%~dp0%SCRIPT_DIR%\extract_and_rename.py"
if not exist "%PYTHON_EXE%" (
echo 配布パッケージが不完全です(pythonフォルダが見つかりません)。管理者に確認してください。
pause
exit /b 1
)
if not exist "%MAIN_PY%" (
echo 配布パッケージが不完全です(extract_and_rename.pyが見つかりません)。管理者に確認してください。
pause
exit /b 1
)
"%PYTHON_EXE%" "%MAIN_PY%"
- Step 2: 手動動作確認(config.txt無しケース)
OCR仕分け実行.bat をリポジトリルート(config.txt がまだ存在しない状態、Task 11で作成予定)でダブルクリックし、「config.txt が見つかりません」メッセージが出て一時停止することを目視確認する。この時点ではまだTask 11未実施のため確認できない場合は、Task 11完了後にまとめて確認してよい。
- Step 3: コミット
git add "OCR仕分け実行.bat"
git commit -m "feat: 起動バッチOCR仕分け実行.batを追加"
Task 11: config.txt雛形・README更新
Files:
-
Create:
config.txt -
Modify:
README.md -
Step 1: config.txt雛形作成
config.txt:
スクリプトフォルダ=スクリプト
テンプレートフォルダ=テンプレート
スキャンフォルダ=スキャン
アウトプットフォルダ=アウトプット
成功フォルダ=成功
失敗フォルダ=失敗
ログフォルダ=ログ
margin=0.10
DPI=300
ファイル安定待ち秒=1
ファイル安定待ちリトライ回数=5
- Step 2: README.md更新
README.md の内容を、旧CLI手順(--template 等の引数説明)から新しい運用手順に置き換える:
# PDFスキャン監視・自動振り分けツール
テンプレート画像(赤枠=氏名欄、青枠=日付欄)を基準に、"スキャン"フォルダ内のPDFから
氏名・日付をOCRで抽出し、リネーム・振り分けを行うツールです。
## 使い方
1. `config.txt` で各フォルダ名・動作パラメータを確認(通常は初期値のまま利用可)
2. `テンプレート/` フォルダにテンプレート画像を1つ入れる(複数ある場合はファイル名が最も早いもの)
- 画像編集ソフト等で、氏名欄に赤い四角枠、日付欄に青い四角枠を重ねて保存(PNG推奨)
- テンプレートは実際の対象PDFと同じ用紙サイズ・向きで作成すること
3. `スキャン/` フォルダに処理対象PDFを入れる
4. `OCR仕分け実行.bat` をダブルクリック
5. 処理結果は以下に振り分けられる:
- 成功: `アウトプット/氏名/氏名_YYYYMMDD.pdf` にコピー、元ファイルは `成功/` へ移動
- 失敗: `失敗/` フォルダへ元ファイルのまま移動
6. 実行ログは `ログ/YYYY-MM-DD.log` に記録される
7. 処理を途中で止めたい場合は、コンソールが表示されている間に何かキーを押す(1件処理する度にチェックされる)
## 開発者向け(配布パッケージのビルド)
配布用の同梱Python・tesseract・popplerは `build/build_package.py` で自動生成する。
詳細は同スクリプトのコメントを参照。ネットワーク接続必須、実行には数分〜数十分かかる。
## 開発者向け(テスト実行)
```bash
python -m pip install pytest
python -m pytest スクリプト/tests -v
注意点
- OCRは完璧ではありません。運用初期は
アウトプット/の結果を必ず目視確認してください - 日付の年が2桁の場合は "20XX年" と仮定して補完します。和暦は非対応です
- 複数ページPDFは1ページ目のみ対象です
- ファイル名衝突時は
氏名_YYYYMMDD(2).pdfのように連番が付与されます
- [ ] **Step 3: コミット**
```bash
git add config.txt README.md
git commit -m "docs: config.txt雛形とREADMEを新運用フローに合わせて更新"
Task 12: build/build_package.py(配布パッケージ自動生成)
Files:
- Create:
build/build_package.py
Interfaces:
- Consumes: リポジトリ内の
OCR仕分け実行.bat、config.txt、スクリプト/*.py - Produces:
dist/ScanOCR/配下の配布用フォルダ一式
このスクリプトはネットワークダウンロード・外部インストーラー実行を伴うため自動テスト対象外とする。実行には7-Zip(7z コマンド)がPATHに通っていることが前提(tesseract-ocrインストーラーの展開に使用)。実際のダウンロード・展開が最後まで成功するかは、この計画のコードレビュー時点では検証できない未知のリスクとして扱う。 Step 3でユーザーが実機で1回実行し、失敗した場合はURLやインストーラーの展開方法を個別に調整する前提とする。
- Step 1: build_package.py実装
build/build_package.py:
# -*- coding: utf-8 -*-
"""
配布パッケージ生成スクリプト。
開発者PC上で1回実行すると、dist/ScanOCR/ 配下に
Python embeddable版・依存pipパッケージ・tesseract-ocr・poppler一式を含む
配布用フォルダを生成する。
前提:
- ネットワーク接続必須
- 7-Zip(7zコマンド)がPATHに通っていること(tesseractインストーラー展開用)
使い方:
python build/build_package.py
"""
from __future__ import annotations
import shutil
import subprocess
import tempfile
import urllib.request
import zipfile
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
DIST_DIR = REPO_ROOT / "dist" / "ScanOCR"
PYTHON_VERSION = "3.11.9"
PYTHON_EMBED_URL = (
f"https://www.python.org/ftp/python/{PYTHON_VERSION}/"
f"python-{PYTHON_VERSION}-embed-amd64.zip"
)
GET_PIP_URL = "https://bootstrap.pypa.io/get-pip.py"
PIP_PACKAGES = ["opencv-python", "pdf2image", "pytesseract", "pillow", "numpy"]
TESSERACT_VERSION = "5.4.0.20240606"
TESSERACT_INSTALLER_URL = (
"https://github.com/UB-Mannheim/tesseract/releases/download/"
f"v{TESSERACT_VERSION}/tesseract-ocr-w64-setup-{TESSERACT_VERSION}.exe"
)
POPPLER_VERSION = "24.02.0-0"
POPPLER_ZIP_URL = (
"https://github.com/oschwartz10612/poppler-windows/releases/download/"
f"v{POPPLER_VERSION}/Release-{POPPLER_VERSION}.zip"
)
CODE_FILES = (
"extract_and_rename.py",
"box_detector.py",
"config_loader.py",
"file_ops.py",
"lock_manager.py",
"logger.py",
)
RUNTIME_FOLDERS = ("テンプレート", "スキャン", "アウトプット", "成功", "失敗", "ログ")
def download(url: str, dest: Path) -> None:
print(f"ダウンロード: {url}")
urllib.request.urlretrieve(url, dest)
def prepare_dist_dir() -> None:
if DIST_DIR.exists():
shutil.rmtree(DIST_DIR)
DIST_DIR.mkdir(parents=True)
def copy_code_files() -> Path:
script_dst = DIST_DIR / "スクリプト"
script_dst.mkdir(parents=True)
for name in CODE_FILES:
shutil.copy2(REPO_ROOT / "スクリプト" / name, script_dst / name)
shutil.copy2(REPO_ROOT / "OCR仕分け実行.bat", DIST_DIR / "OCR仕分け実行.bat")
shutil.copy2(REPO_ROOT / "config.txt", DIST_DIR / "config.txt")
return script_dst
def setup_python(script_dst: Path, tmp_dir: Path) -> None:
python_dir = script_dst / "python"
python_dir.mkdir(parents=True)
zip_path = tmp_dir / "python-embed.zip"
download(PYTHON_EMBED_URL, zip_path)
with zipfile.ZipFile(zip_path) as zf:
zf.extractall(python_dir)
pth_files = list(python_dir.glob("python*._pth"))
if not pth_files:
raise RuntimeError("python*._pth が見つかりません(embeddable zipの構成が変わった可能性)")
pth_path = pth_files[0]
content = pth_path.read_text(encoding="utf-8")
content = content.replace("#import site", "import site")
pth_path.write_text(content, encoding="utf-8")
get_pip_path = tmp_dir / "get-pip.py"
download(GET_PIP_URL, get_pip_path)
python_exe = python_dir / "python.exe"
subprocess.run([str(python_exe), str(get_pip_path)], check=True)
subprocess.run(
[str(python_exe), "-m", "pip", "install", *PIP_PACKAGES],
check=True,
)
def setup_tesseract(script_dst: Path, tmp_dir: Path) -> None:
tools_dst = script_dst / "tools" / "tesseract"
tools_dst.mkdir(parents=True)
installer_path = tmp_dir / "tesseract-setup.exe"
download(TESSERACT_INSTALLER_URL, installer_path)
extract_dir = tmp_dir / "tesseract-extract"
extract_dir.mkdir()
subprocess.run(
["7z", "x", str(installer_path), f"-o{extract_dir}", "-y"],
check=True,
)
for item in extract_dir.iterdir():
shutil.move(str(item), str(tools_dst / item.name))
def setup_poppler(script_dst: Path, tmp_dir: Path) -> None:
tools_parent = script_dst / "tools"
tools_parent.mkdir(parents=True, exist_ok=True)
tools_dst = tools_parent / "poppler"
zip_path = tmp_dir / "poppler.zip"
download(POPPLER_ZIP_URL, zip_path)
extract_dir = tmp_dir / "poppler-extract"
with zipfile.ZipFile(zip_path) as zf:
zf.extractall(extract_dir)
inner_items = list(extract_dir.iterdir())
source_dir = inner_items[0] if len(inner_items) == 1 else extract_dir
shutil.move(str(source_dir), str(tools_dst))
def create_runtime_folders() -> None:
for name in RUNTIME_FOLDERS:
(DIST_DIR / name).mkdir(parents=True, exist_ok=True)
def main() -> None:
prepare_dist_dir()
script_dst = copy_code_files()
with tempfile.TemporaryDirectory() as tmp:
tmp_dir = Path(tmp)
setup_python(script_dst, tmp_dir)
setup_tesseract(script_dst, tmp_dir)
setup_poppler(script_dst, tmp_dir)
create_runtime_folders()
print(f"配布パッケージ生成完了: {DIST_DIR}")
if __name__ == "__main__":
main()
- Step 2: .gitignoreにdist/が含まれていることを確認
Task 1で作成した .gitignore に既に dist/ が含まれているため追加作業は不要。git status で dist/ がトラッキング対象外になっていることを確認する。
Run: git status --short
Expected: dist/ 配下のファイルが(生成されていても)一覧に出ない
- Step 3: コミット
git add build/build_package.py
git commit -m "feat: 配布パッケージ自動生成スクリプトを追加"
注記(ユーザー向け確認事項): このタスクはコードとして実装するが、実際に最後まで実行してdist/ScanOCR/配下が正しく動く状態になるかは、7-Zipインストール状況・各ダウンロードURLの生存・tesseractインストーラーの展開可否に依存する。初回実行時にエラーが出た場合は、エラーメッセージを元に該当関数(setup_tesseractが最もリスクが高い)を調整する。
全体自己レビュー結果
- spec網羅性: 設計書の全セクション(配布パッケージ構成/config.txt仕様/起動フロー/メインループ/終了処理/ログ仕様/ビルドスクリプト/対象外事項)に対応するタスクを配置済み
- placeholder: なし。全ステップに実コードを記載
- 型・シグネチャ一貫性:
run_queue・copy_with_unique_name・move_with_unique_name・append_log・acquire_lockの引数名はTask間で統一済み - 既知の未実装事項:
config.txtのDPI項目は雛形には存在するがprocess_pdfには未接続(Task 9 Step 3に明記)。将来対応が必要ならフォローアップタスクとする