Xây dựng một hệ thống crawl tài liệu từ intra-mart platform
Find a file
Do Duy eb46e739e7 Add intra-mart documentation crawler
Playwright-based crawler that renders pages from api.intra-mart.jp and
document.intra-mart.jp, extracts main content, and converts to Markdown
via Pandoc. SQLite-backed queue drives a resumable sequential pipeline
across crawler, extractor, converter, and storage modules.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 16:29:54 +07:00
converter Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
crawler Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
extractor Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
models Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
storage Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
utils Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
.gitignore Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
.python-version Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
CLAUDE.md Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
config.yaml Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
LICENSE Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
main.py Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
pyproject.toml Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
README.md Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
requirements.txt Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00
uv.lock Add intra-mart documentation crawler 2026-05-20 16:29:54 +07:00

Intra-Mart Documentation Crawler → Markdown Exporter

🚀 Quickstart với uv (khuyến nghị)

Project đã được lock vào uv.lock + ghim Python qua .python-version. uv sẽ tự cài đúng phiên bản Python, dựng .venv, và sync deps từ lock file — không cần pip, không cần pyenv.

1. Cài uv

OS Lệnh
macOS / Linux curl -LsSf https://astral.sh/uv/install.sh | sh
Windows (PowerShell) powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
macOS (Homebrew) brew install uv
Bất kỳ OS nào có pipx pipx install uv

Xác nhận đã cài: uv --version.

2. Cài Pandoc (bắt buộc, không nằm trong uv)

MarkdownConverter gọi binary pandoc qua subprocess, nên Pandoc phải có sẵn trên PATH.

OS Lệnh
macOS (Homebrew) brew install pandoc
Ubuntu / Debian sudo apt install pandoc
Fedora / RHEL sudo dnf install pandoc
Arch sudo pacman -S pandoc
Windows (winget) winget install --id JohnMacFarlane.Pandoc
Windows (Chocolatey) choco install pandoc
Khác Tải installer tại https://pandoc.org/installing.html

Xác nhận: pandoc --version.

3. Clone & sync deps

git clone https://github.com/IricsDo/intra_mart_doc_crawler.git
cd intra_mart_doc_crawler

uv sync                          # tạo .venv, cài đúng Python 3.11 + toàn bộ deps từ uv.lock
uv run playwright install chromium   # tải Chromium cho Playwright

Trên Linux, có thể playwright báo thiếu system lib (libnss3, libatk, …). Fix: uv run playwright install-deps chromium (cần sudo, dùng apt). Trên macOS/Windows thường không cần.

4. Chạy crawler

uv run python main.py

Hoặc activate venv rồi chạy như Python thường:

# macOS / Linux
source .venv/bin/activate
python main.py

# Windows (PowerShell)
.venv\Scripts\Activate.ps1
python main.py

# Windows (cmd.exe)
.venv\Scripts\activate.bat
python main.py

Chạy lại = resume từ SQLite queue (data/metadata.db). URL đã done bị skip.

5. Dev tools (lint / format / type-check / test)

uv sync --extra dev          # cài thêm black, ruff, mypy, pytest

uv run ruff check .
uv run black .
uv run mypy .
uv run pytest                # asyncio_mode = auto đã preset trong pyproject.toml

Lệnh uv thường dùng

Tác dụng Lệnh
Sync env đúng theo uv.lock uv sync
Sync + cài optional dep dev uv sync --extra dev
Thêm 1 package mới uv add <package>
Bỏ 1 package uv remove <package>
Update lock theo pyproject.toml uv lock
Upgrade 1 package uv lock --upgrade-package <package>
Chạy lệnh trong venv mà không activate uv run <command>

Vẫn dùng pip được — xem mục Cài đặt môi trường (pip) ở dưới. Nhưng uv nhanh hơn 10100× và đảm bảo cùng phiên bản deps trên mọi máy nhờ uv.lock.


Mục tiêu project

Crawl tài liệu từ:

rồi convert toàn bộ sang Markdown có:

  • giữ nguyên hierarchy theo URL path

  • giữ internal links (rewrite sang đường dẫn local)

  • tải assets (ảnh, css, file đính kèm) về local

  • hỗ trợ JS-rendered docs (Playwright + Chromium)

  • resume được khi chạy lại (state nằm trong SQLite)

  • output sạch để dùng với:

    • Obsidian
    • MkDocs
    • Docusaurus
    • RAG / LLM
    • GitBook

Kiến trúc tổng thể

Toàn bộ pipeline nằm trong main.py (DocumentationCrawler), điều phối các module con. Flow mỗi page:

URLManager.normalize / is_allowed
        ↓
PlaywrightClient.fetch_page          (Chromium, JS-rendered, retry qua tenacity)
        ↓
HTMLCleaner.clean                    (loại REMOVE_SELECTORS)
        ↓
HTMLExtractor.extract_main_content   (match CONTENT_SELECTORS đầu tiên)
        ↓
AssetDownloader                      (download + rewrite asset link sang local)
        ↓
MarkdownConverter.html_to_markdown   (Pandoc → GFM, kèm YAML frontmatter)
        ↓
FileWriter / PathMapper              (output/markdown/<domain>/<sanitized path>.md)
        ↓
extract_links + SphinxDiscovery + NavigationParser → enqueue URL mới

State nằm ở SQLite, không phải in-memory

storage/metadata_db.py (file data/metadata.db) là source of truth.

Bảng chính:

  • crawl_queue — drive vòng lặp crawl: pending / processing / retry / failed / done. Lấy URL kế tiếp theo priority DESC, depth ASC, discovered_at ASC.
  • pages — metadata page crawl thành công + đường dẫn markdown output.
  • assets, page_links, discovery_sources, failures — provenance + site graph.

Chạy lại python main.py sẽ resume: URL done bị bỏ qua, URL đang processing bị reset về retry khi khởi động.

⚠️ Crawl loop hiện chạy sequential single-worker (chưa có asyncio worker pool). Để giới hạn memory, Playwright được restart mỗi BROWSER_RECYCLE_EVERY = 500 page.

Discovery đa nguồn

URL được phát hiện từ 3 nguồn, merge lại trước khi enqueue:

  1. DocumentationCrawler.extract_links — generic <a href> + selector đặc thù Sphinx (a.reference.internal, .toctree-wrapper a, link[rel=next/prev/up]).
  2. SphinxDiscovery — probe searchindex.js, genindex.html, …
  3. NavigationParser — dựng nav tree để export SUMMARY.md / navigation_tree.json.

URL → filesystem mapping

storage/path_mapper.py map sang <output_dir>/<domain>/<sanitized path>.md:

  • sanitize ký tự Windows-invalid + reserved names
  • hash query string thành suffix __<hash>
  • rút gọn tên > 180 ký tự bằng MD5 tail

Asset link rewrite trong extractor/asset_downloader.py tính path tương đối với output path của page hiện tại — nếu đổi layout output của PathMapper, phải sửa luôn rewrite_asset_links.


Công nghệ sử dụng

Thành phần Công nghệ
Rendering Playwright (Chromium)
HTTP fast-path httpx
HTML Parse BeautifulSoup4 + lxml
URL handling urllib.parse
Markdown conversion Pandoc (qua pypandoc / markdownify)
State & queue SQLite
Logging loguru
Retry tenacity
Config constants ở đầu main.py (xem mục Configuration)

Cấu trúc project

intra-mart-crawler/
│
├── README.md
├── CLAUDE.md
├── requirements.txt
├── pyproject.toml
├── config.yaml                ← KHÔNG được main.py đọc (xem Configuration)
├── main.py                    ← entrypoint + DocumentationCrawler
│
├── crawler/
│   ├── crawler.py
│   ├── http_client.py
│   ├── playwright_client.py
│   ├── url_manager.py
│   ├── sitemap_parser.py
│   ├── sphinx_discovery.py
│   └── navigation_parser.py
│
├── extractor/
│   ├── html_extractor.py
│   ├── html_cleaner.py
│   └── asset_downloader.py
│
├── converter/
│   ├── markdown_converter.py
│   └── link_rewriter.py
│
├── storage/
│   ├── file_writer.py
│   ├── metadata_db.py         ← SQLite source of truth
│   └── path_mapper.py
│
├── models/
│   ├── page.py
│   └── crawl_result.py
│
├── utils/
│   ├── logger.py
│   ├── retry.py
│   └── helpers.py
│
├── output/                    ← runtime artifact (gitignored)
│   ├── markdown/
│   └── assets/
│
├── data/                      ← runtime artifact (gitignored)
│   └── metadata.db
│
└── logs/                      ← rotating logs (gitignored)
    ├── crawler.log
    ├── errors.log
    └── debug.log

Cài đặt môi trường (pip — cách cũ)

Nếu bạn dùng uv thì có thể bỏ qua section này — xem Quickstart với uv ở đầu file.

Yêu cầu: Python >= 3.11Pandoc có trên PATH.

# 1. Python deps
pip install -r requirements.txt

# 2. Browser cho Playwright
playwright install chromium

# 3. Pandoc
sudo apt install pandoc          # Ubuntu/Debian
# macOS:   brew install pandoc
# Windows: winget install --id JohnMacFarlane.Pandoc
#          hoặc https://pandoc.org/installing.html

Dev environment (lint / format / type-check / test)

pip install -e ".[dev]"

ruff check .
black .
mypy .
pytest                                  # asyncio_mode = auto đã preset
pytest tests/test_x.py::test_name       # chạy 1 test

Line length pinned ở 79 cho cả blackruff. Target Python 3.11.


Configuration

⚠️ config.yaml không được main.py đọc. Đây là tàn dư từ thiết kế cũ — sẽ refactor sau.

Cấu hình "live" là các hằng số ở đầu main.py:

Constant Ý nghĩa
START_URLS Seed URLs để crawl
ALLOWED_DOMAINS Whitelist domain — ngoài list này sẽ không enqueue
OUTPUT_DIR Thư mục output gốc (mặc định output/)
REMOVE_SELECTORS CSS selectors bị drop trước khi extract
CONTENT_SELECTORS Selectors thử lần lượt để lấy main content
MAX_PAGES_PER_RUN Số page tối đa mỗi lần chạy
MAX_ATTEMPTS Số lần retry trước khi mark failed
BROWSER_RECYCLE_EVERY Restart Chromium sau mỗi N page (giới hạn memory leak)

Muốn đổi behaviour → sửa các constant này, không phải config.yaml.


Chạy project

python main.py

Chạy lại lệnh trên = resume từ SQLite queue. URL đã done bị skip, URL processing reset về retry.


Output ví dụ

output/
├── markdown/
│   ├── api.intra-mart.jp/
│   │   └── iap/
│   │       ├── index.md
│   │       ├── javadoc/
│   │       │   └── ...
│   │       └── apilist-ssjs/
│   │           └── ...
│   │
│   └── document.intra-mart.jp/
│       └── library/
│           ├── getting-started.md
│           └── workflow.md
│
└── assets/
    └── ...    (ảnh, css, file đính kèm)

Mỗi file .md có YAML frontmatter (source URL, crawl time, ...) và đã rewrite internal link + asset link sang relative path.

Ngoài ra:

  • SUMMARY.md — index theo nav tree (cho GitBook / mdBook).
  • navigation_tree.json — nav tree dạng JSON (cho MkDocs / Docusaurus / custom sidebar).

Logging

utils/logger.py cấu hình loguru ngay khi import (chỉ cần from utils.logger import logger). Ghi rotating log vào logs/crawler.log, logs/errors.log, logs/debug.log — không cần setup gì thêm.


Hướng phát triển tiếp theo

Trong scope crawler

  1. Async concurrent crawling — hiện crawl sequential, có thể thêm asyncio worker pool + semaphore + rate limiting.
  2. Better extraction fallback — bổ sung trafilatura / readability-lxml cho page không match CONTENT_SELECTORS nào.
  3. Duplicate detection — hash HTML/markdown content để skip page trùng nội dung khác URL.
  4. Đọc config.yaml thực sự (hoặc bỏ hẳn file này) — đồng bộ với constants trong main.py.
  5. Incremental update — re-crawl chỉ những page có ETag / Last-Modified thay đổi.

Hệ sinh thái sau crawler

  1. Build MkDocs / Docusaurus site từ output.
  2. Build vector database + embedding docs.
  3. Tạo RAG chatbot trên docs đã embed.
  4. Semantic search UI.
  5. Git sync pipeline cho output markdown.