fix: resolve AWS WAF challenge, add URL sanitization, robust DOM/JSON-LD fallback extraction, and update README
This commit is contained in:
54
README.md
54
README.md
@@ -0,0 +1,54 @@
|
||||
# 🕵️ Scraper Service
|
||||
|
||||
A robust background scraping service built with **Playwright**, **BullMQ**, and **Chromium** to scrape hotel metadata, real-time prices, Genius discounts, and guest reviews from Booking.com.
|
||||
|
||||
---
|
||||
|
||||
## 🌟 Key Capabilities
|
||||
|
||||
- **Deep Schema & Metadata Extraction:**
|
||||
- Unpacks nested Schema.org JSON-LD scripts (including `@graph` arrays and accommodation schemas: `Hotel`, `Resort`, `Apartment`, `Hostel`, `LodgingBusiness`, etc.).
|
||||
- Multi-tier DOM and OpenGraph meta tag fallbacks for title, address, hero images, and ratings.
|
||||
- Automatic URL slug parser (`extractHotelNameFromUrl`) as an ultimate fallback if client-side rendering is obstructed.
|
||||
- **AWS WAF Challenge & Anti-Bot Evasion:**
|
||||
- **URL Sanitization (`cleanBookingUrl`):** Automatically cleans tracking tokens and expired challenge parameters (`chal_t`, `force_referer`, `sid`, `srpvid`, `srepoch`).
|
||||
- **Browser Fingerprint Masking:** Emulates standard desktop Chrome headers, `window.chrome` runtime objects, `navigator.plugins`, and German language headers.
|
||||
- **Active WAF Challenge Resolution:** Automatically detects AWS WAF challenge pages and waits for client-side JavaScript token generation and redirection.
|
||||
- **Overlay & Banner Handling:**
|
||||
- Auto-dismisses OneTrust cookie banners, modal dialogs, and selection popups before evaluating DOM data.
|
||||
- **Dual-Mode Scraping:**
|
||||
- **Quick Mode (Stammdaten & Prices):** Scrapes name, star rating, review count, address, hero image, room size, and prices in ~3–5 seconds and posts directly to Next.js webhook without LLM overhead.
|
||||
- **Full AI Review Mode:** Sorts reviews by newest first, pages through reviews, filters for western travelers from the last 2 years, pulls Google Maps bathroom photos via SerpApi, and sends reviews to n8n for LLM summarization.
|
||||
- **Debugging & Inspection:**
|
||||
- Automatically captures post-render rendered DOM (`src/hotel-debug.html`) and screenshots (`src/hotel-debug.png`) when `DEBUG_HTML=true` or when requested via `?debug=true`.
|
||||
|
||||
---
|
||||
|
||||
## ⚙️ Configuration (`.env` / Docker Environment)
|
||||
|
||||
| Variable | Description | Default |
|
||||
|---|---|---|
|
||||
| `REDIS_HOST` | Redis host or container name | `redis` |
|
||||
| `REDIS_PORT` | Redis port | `6379` |
|
||||
| `DEBUG_HTML` | Enable saving `hotel-debug.html` and `hotel-debug.png` | `true` |
|
||||
| `SERPAPI_KEY` | (Optional) SerpApi key for Google Maps bathroom photos | `null` |
|
||||
|
||||
---
|
||||
|
||||
## 🚀 Running with Docker Compose
|
||||
|
||||
```bash
|
||||
docker compose up -d --build
|
||||
```
|
||||
|
||||
### Viewing Logs
|
||||
```bash
|
||||
docker logs -f scraper-app
|
||||
```
|
||||
|
||||
### Inspecting Scraped DOM Debug Files
|
||||
When debugging hotel pages, the rendered DOM and screenshot are saved directly to:
|
||||
- `src/hotel-debug.html`
|
||||
- `src/hotel-debug.png`
|
||||
- `src/review-debug.png`
|
||||
|
||||
|
||||
Reference in New Issue
Block a user