# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Overview Documentation crawler that renders pages from `api.intra-mart.jp` and `document.intra-mart.jp` with Playwright, extracts the main content, and converts it to Markdown via Pandoc. Output is structured for Obsidian / MkDocs / Docusaurus / RAG ingestion. ## Commands ```bash # Install (Python >= 3.11) pip install -r requirements.txt playwright install chromium # pandoc must be on PATH (apt install pandoc / https://pandoc.org/installing.html) # Run the crawler python main.py # Lint / format / typecheck / test (dev extras) pip install -e ".[dev]" ruff check . black . mypy . pytest # asyncio_mode=auto is preset pytest tests/test_x.py::test_name # single test ``` Re-running `python main.py` resumes from the SQLite queue — finished URLs are skipped, `processing` rows are reset to `retry` on startup. ## Architecture The pipeline lives in `main.py` as a single `DocumentationCrawler` class that orchestrates per-module components. Flow per page: ``` URLManager.normalize/is_allowed ↓ PlaywrightClient.fetch_page (Chromium, JS-rendered, retries via tenacity) ↓ HTMLCleaner.clean (drop REMOVE_SELECTORS) ↓ HTMLExtractor.extract_main_content (first match of CONTENT_SELECTORS) ↓ AssetDownloader (download + rewrite asset links to local paths) ↓ MarkdownConverter.html_to_markdown (Pandoc → GFM, with YAML frontmatter) ↓ FileWriter / PathMapper (output/markdown//.md) ↓ extract_links + SphinxDiscovery + NavigationParser → enqueue new URLs ``` ### State is in SQLite, not in memory `storage/metadata_db.py` (`data/metadata.db`) is the source of truth. Tables: - `crawl_queue` — drives the loop (`pending` / `processing` / `retry` / `failed` / `done`); `get_next_pending_url` orders by `priority DESC, depth ASC, discovered_at ASC`. - `pages` — successful crawl metadata + markdown output path. - `assets`, `page_links`, `discovery_sources`, `failures` — provenance and the site graph. The crawl loop is **sequential single-worker** even though README mentions concurrency — there is no asyncio worker pool. `BROWSER_RECYCLE_EVERY = 500` restarts Playwright periodically to bound memory. ### Discovery is multi-source Links are discovered three ways and merged before enqueue: (1) generic `` extraction in `DocumentationCrawler.extract_links` (also targets Sphinx-specific selectors like `a.reference.internal`, `.toctree-wrapper a`, `link[rel=next/prev/up]`), (2) `SphinxDiscovery` (probes `searchindex.js`, `genindex.html`, etc.), (3) `NavigationParser` (builds a nav tree for `SUMMARY.md` / `navigation_tree.json` export). ### URL → filesystem mapping `storage/path_mapper.py` produces `//.md`. Sanitizes Windows-invalid chars + reserved names, hashes query strings into a `__` suffix, and shortens names > 180 chars with an MD5 tail. Asset link rewriting in `extractor/asset_downloader.py` computes paths *relative to the current page's output path*, so don't change `PathMapper` output layout without also revisiting `rewrite_asset_links`. ### Configuration `config.yaml` exists but **`main.py` does not read it** — the live configuration is the module-level constants at the top of `main.py` (`START_URLS`, `ALLOWED_DOMAINS`, `REMOVE_SELECTORS`, `CONTENT_SELECTORS`, `MAX_PAGES_PER_RUN`, `BROWSER_RECYCLE_EVERY`, `MAX_ATTEMPTS`). Update those, not the YAML. ### Logging `utils/logger.py` configures loguru on import (just `from utils.logger import logger`). Writes rotating logs to `logs/crawler.log`, `logs/errors.log`, `logs/debug.log` — no extra setup needed. ## Conventions - Line length 79 (`black` and `ruff` both pinned to this). Target Python 3.11. - `pytest-asyncio` runs in `auto` mode — async test functions need no decorator. - Output dirs (`output/`, `data/`, `logs/`) are runtime artifacts; `output/assets/`, `output/markdown/`, `data/` are gitignored.