# Audio Transcription Desktop Milestone 1 Implementation Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Deliver a Windows desktop GUI that selects and inspects media, manages secure settings, and exposes a clearly stubbed transcription action. **Architecture:** A `src/voice_transcriptor` package separates typed models and services from PySide6 UI modules. FFprobe and settings logic are synchronous, testable services; the UI runs media probing in a Qt thread pool and ignores stale results. **Tech Stack:** Python 3.12+, PySide6, keyring, OpenAI Python SDK, pytest, pytest-qt, FFmpeg/ffprobe. **Spec:** `docs/superpowers/specs/2026-08-29-transcription-desktop-milestone-1-design.md` ## Global Constraints - Target Windows 11 and Python 3.12 or newer. - Store API keys only through `keyring`; never serialize or log them. - Default language is `pt-BR`. - Transcription remains stubbed in milestone 1. - Keep UI imports out of business services and retain PyInstaller-compatible entry points. - Preserve the empty tracked `README.md` by filling it with project documentation. --- ## File Map - `pyproject.toml`: package metadata, entry point, and pytest settings. - `requirements.txt`: runtime and test dependencies for straightforward Windows setup. - `.gitignore`: Python, test, environment, editor, and PyInstaller artifacts. - `src/voice_transcriptor/models.py`: immutable shared data values. - `src/voice_transcriptor/formatting.py`: file-size and duration formatting. - `src/voice_transcriptor/services/media_probe.py`: tool discovery, file validation, ffprobe execution and parsing. - `src/voice_transcriptor/services/settings.py`: defaults and atomic non-secret JSON persistence. - `src/voice_transcriptor/services/credentials.py`: keyring-only API-key access. - `src/voice_transcriptor/services/transcription.py`: explicit milestone stub. - `src/voice_transcriptor/ui/settings_dialog.py`: settings form and validation. - `src/voice_transcriptor/ui/main_window.py`: selection/drop UI, async probing, metadata and status presentation. - `src/voice_transcriptor/app.py`, `__main__.py`: dependency wiring and executable entry points. - `tests/`: focused unit tests and an offscreen Qt smoke test. - `README.md`: Windows setup, execution, security, testing, and milestone limitations. ### Task 1: Typed core, formatting, and project scaffolding **Files:** - Create: `pyproject.toml`, `requirements.txt`, `.gitignore` - Create: `src/voice_transcriptor/__init__.py`, `src/voice_transcriptor/models.py`, `src/voice_transcriptor/formatting.py` - Test: `tests/test_formatting.py` **Interfaces:** - Produces: `AppSettings(model: str, language: str, output_directory: Path)`, `ToolStatus(ffmpeg_path: Path | None, ffprobe_path: Path | None)`, `MediaInfo(path: Path, size_bytes: int, duration_seconds: float | None, audio_codec: str | None)`, `format_file_size(int) -> str`, and `format_duration(float | None) -> str`. - [ ] **Step 1: Write failing formatting tests** for byte units, multi-hour duration, and unavailable duration. - [ ] **Step 2: Run `python -m pytest tests/test_formatting.py -v`** and confirm collection/import failure. - [ ] **Step 3: Add package configuration, typed dataclasses, and minimal pure formatting functions.** Size uses binary units and duration renders `HH:MM:SS`; unavailable duration renders `Unknown`. - [ ] **Step 4: Run `python -m pytest tests/test_formatting.py -v`** and confirm all cases pass. - [ ] **Step 5: Commit with `git commit -m "build: scaffold typed application core"`.** ### Task 2: Media inspection service **Files:** - Create: `src/voice_transcriptor/services/__init__.py`, `src/voice_transcriptor/services/media_probe.py` - Test: `tests/test_media_probe.py` **Interfaces:** - Consumes: `MediaInfo`, `ToolStatus`. - Produces: `detect_tools() -> ToolStatus`, `validate_media_path(path: Path) -> Path`, `parse_probe_output(path: Path, size_bytes: int, payload: str) -> MediaInfo`, `MediaProbeService.probe(path: Path) -> MediaInfo`, and typed `MediaProbeError` subclasses. - [ ] **Step 1: Write failing tests** for executable discovery, path rejection, audio-only and video-with-audio JSON, missing fields, malformed JSON, nonzero subprocess result, and timeout mapping. - [ ] **Step 2: Run `python -m pytest tests/test_media_probe.py -v`** and confirm import failure. - [ ] **Step 3: Implement safe probing** with `subprocess.run([...], capture_output=True, text=True, timeout=30, check=False, creationflags=CREATE_NO_WINDOW on Windows)` and no shell. - [ ] **Step 4: Run the focused tests** and confirm they pass. - [ ] **Step 5: Commit with `git commit -m "feat: add ffprobe media inspection"`.** ### Task 3: Settings and credential persistence **Files:** - Create: `src/voice_transcriptor/services/settings.py`, `src/voice_transcriptor/services/credentials.py` - Test: `tests/test_settings.py`, `tests/test_credentials.py` **Interfaces:** - Consumes: `AppSettings`. - Produces: `SettingsRepository(path: Path | None = None)`, `.load() -> tuple[AppSettings, str | None]`, `.save(settings: AppSettings) -> None`; `CredentialService(backend=keyring)`, `.get_api_key() -> str | None`, `.has_api_key() -> bool`, `.set_api_key(value: str) -> None`; `SettingsError`, `CredentialError`. - [ ] **Step 1: Write failing tests** for defaults, round trip, malformed recovery warning, atomic replacement, absence of API-key fields, and mocked keyring calls/failures. - [ ] **Step 2: Run both focused test files** and confirm import failure. - [ ] **Step 3: Implement per-user configuration** under `%APPDATA%/VoiceTranscriptor/settings.json`, using UTF-8 JSON and `Path.replace` from a sibling temporary file. - [ ] **Step 4: Implement the keyring adapter** with stable service `voice-transcriptor` and account `openai-api-key`; reject blank keys and sanitize exceptions. - [ ] **Step 5: Run focused tests** and confirm they pass. - [ ] **Step 6: Commit with `git commit -m "feat: persist settings and credentials securely"`.** ### Task 4: Transcription boundary **Files:** - Create: `src/voice_transcriptor/services/transcription.py` - Test: `tests/test_transcription.py` **Interfaces:** - Produces: `TranscriptionNotImplementedError` and `TranscriptionService.transcribe(media: MediaInfo, settings: AppSettings) -> None`. - [ ] **Step 1: Write a failing test** asserting the service raises the specific milestone exception with a user-readable message. - [ ] **Step 2: Run the focused test** and confirm import failure. - [ ] **Step 3: Implement the minimal typed stub.** - [ ] **Step 4: Run the focused test** and confirm it passes. - [ ] **Step 5: Commit with `git commit -m "feat: define transcription service boundary"`.** ### Task 5: Settings dialog **Files:** - Create: `src/voice_transcriptor/ui/__init__.py`, `src/voice_transcriptor/ui/settings_dialog.py` - Test: `tests/ui/test_settings_dialog.py` **Interfaces:** - Consumes: `AppSettings`, `SettingsRepository`, `CredentialService`. - Produces: `SettingsDialog(settings, repository, credentials, parent=None)` and `settings_saved(AppSettings)` signal. - [ ] **Step 1: Write offscreen Qt tests** for populated defaults, masked key entry, preserving an existing credential when blank, saving a supplied key, output-directory validation, and no secret in widget/log output after save. - [ ] **Step 2: Run the focused UI tests** and confirm import failure. - [ ] **Step 3: Build a native `QDialog` form** with editable model/language, folder picker, password echo mode, Save/Cancel buttons, and concise message boxes on service failure. - [ ] **Step 4: Run the focused tests** and confirm they pass. - [ ] **Step 5: Commit with `git commit -m "feat: add secure settings dialog"`.** ### Task 6: Main window and asynchronous probe flow **Files:** - Create: `src/voice_transcriptor/ui/main_window.py` - Test: `tests/ui/test_main_window.py` **Interfaces:** - Consumes all service interfaces, formatting helpers, and shared models. - Produces: `MainWindow(media_probe, settings_repository, credentials, transcription_service)` with `select_file(path: Path) -> None`; internal `ProbeSignals` and `ProbeWorker(QRunnable)` communicate results without blocking. - [ ] **Step 1: Write offscreen Qt tests** for the visible controls, supported single-file drops, rejected URLs/multiple files, populated metadata, missing-tool status, stale probe suppression, settings opening, validation on Transcribe, and stub status logging. - [ ] **Step 2: Run focused UI tests** and confirm import failure. - [ ] **Step 3: Implement the native main-window layout** with a drop frame, Browse and Settings buttons, four metadata rows, tool status, prominent Transcribe button, and read-only log. - [ ] **Step 4: Implement QRunnable probing** tagged with monotonically increasing request IDs; callbacks update only when the ID matches the current selection. - [ ] **Step 5: Implement button actions and readable error handling** without logging credentials or full tracebacks in the GUI. - [ ] **Step 6: Run focused UI tests** and confirm they pass. - [ ] **Step 7: Commit with `git commit -m "feat: add media selection desktop interface"`.** ### Task 7: Application wiring and startup smoke test **Files:** - Create: `src/voice_transcriptor/app.py`, `src/voice_transcriptor/__main__.py` - Test: `tests/ui/test_app.py` **Interfaces:** - Produces: `create_main_window() -> MainWindow` and `main() -> int`; console entry point `voice-transcriptor`. - [ ] **Step 1: Write a failing offscreen smoke test** that constructs, shows, processes events for, and closes the window. - [ ] **Step 2: Run the focused test** and confirm import failure. - [ ] **Step 3: Wire concrete services and add guarded startup error reporting.** Reuse an existing `QApplication` in tests and create one in `main()`. - [ ] **Step 4: Run the smoke test and `python -m compileall -q src tests`.** - [ ] **Step 5: Commit with `git commit -m "feat: wire desktop application startup"`.** ### Task 8: Documentation and full validation **Files:** - Modify: `README.md`, `requirements.txt` **Interfaces:** - Documents the `python -m voice_transcriptor` and `voice-transcriptor` entry points and environment setup. - [ ] **Step 1: Write README instructions** for Python 3.12 virtual environments, dependency installation, FFmpeg installation/PATH checks, running, testing, Windows Credential Manager security, known milestone limitations, and future PyInstaller packaging. - [ ] **Step 2: Install dependencies with `python -m pip install -r requirements.txt`** when they are not already available. - [ ] **Step 3: Run `python -m pytest -v` and `python -m compileall -q src tests`.** - [ ] **Step 4: Run `ffmpeg -version` and `ffprobe -version`** and record availability without treating absence as a test failure. - [ ] **Step 5: Launch an offscreen controlled startup** that opens the main window, processes events, and closes it via a timer. - [ ] **Step 6: Review `git diff --check`, `git status --short`, and scan tracked project text for accidental API-key-like values.** - [ ] **Step 7: Commit with `git commit -m "docs: add Windows setup and validation guide"`.**