Files

22 lines
2.0 KiB
Markdown

# Engineering & Scraping Guidelines
## 1. Git Commit & Repository Hygiene
- **Conventional Commits:** Write clear, concise commit messages prefixed with standard types (e.g., `feat:`, `fix:`, `refactor:`, `docs:`).
- **Exclude Transient Artifacts:** Never stage or commit generated debug files (e.g., `src/*.html`, `src/*.png`, large logs) to version control. Always maintain `.gitignore` to keep git packs lean and prevent remote unpacker errors on Gitea/GitHub.
## 2. Web Scraping & DOM Inspection Strategy
- **Rendered DOM Dumping for Debugging:** When debugging extraction failures on dynamic or anti-bot protected pages:
- Capture the post-render HTML via `await page.content()` and save it to `src/hotel-debug.html`.
- Capture a synchronized screenshot (`src/hotel-debug.png`).
- Inspect the raw rendered DOM to verify if the page was served as standard SSR, an A/B test SPA variant, or intercepted by a challenge (e.g. AWS WAF).
- **Multi-Tier Extraction Hierarchy:** Always implement fallbacks in this order:
1. **Deep JSON-LD Parsing:** Traverse all `<script type="application/ld+json">` tags, including `@graph` arrays and varied accommodation schemas (`Hotel`, `Resort`, `Apartment`, `Hostel`, `LodgingBusiness`, etc.).
2. **DOM Selectors:** Fallback to standard test IDs, classes, and element hierarchy (`data-testid`, header titles, score components).
3. **Meta Tags:** Fallback to OpenGraph and Twitter meta tags (`og:title`, `og:image`, `og:description`).
4. **URL Slug Extraction:** As an ultimate safety net, derive clean human-readable names from the URL path if the HTML body was obstructed.
- **Anti-Bot & WAF Handling:**
- Sanitize input URLs by stripping transient tokens (`chal_t`, `force_referer`, session IDs).
- Emulate full browser properties (`window.chrome`, `navigator.plugins`, `navigator.languages`).
- Actively detect WAF challenge interstitial pages and wait for client-side JavaScript token computation and automatic redirection before evaluating the DOM.