# Agent Notes ## Bot detection (blank page symptom) Goodreads returns `` to headless browsers it identifies as automated. Symptoms: - `page.content()` → bare empty HTML skeleton - `page.title()` → empty string - Screenshot is fully white - `.exportBooks` div count is 0 **Root cause:** Playwright's default headless Chromium exposes `navigator.webdriver = true` and announces itself via the `AutomationControlled` Blink feature. **Fix applied:** - Launch with `--disable-blink-features=AutomationControlled` - Set a realistic `user_agent` and `viewport` on the context - `context.add_init_script(...)` to set `navigator.webdriver = undefined` - Guard the `logged_in` check: a blank page must not be treated as "logged in" ## Goodreads Import/Export page (`/review/import`) The "Export Library" button is present in the static HTML as: ```html ``` ### Why `get_by_role("button", name="Export Library")` can fail `page.goto()` only waits for the `load` event. JavaScript may still be running (React hydration, etc.) when the button is queried, so `count()` can return 0 or `is_visible()` can return `False` even though the button exists in the DOM. **Fix applied:** call `page.wait_for_load_state("networkidle")` after `goto`, then use `wait_for(state="visible", timeout=10s)` instead of `is_visible()`. A CSS-class fallback (`button.js-LibraryExport`) is also tried in case the accessible-name lookup fails. ### Export flow 1. Click "Export Library" → Goodreads POSTs to `/review_porter/export/` 2. The `#exportFile` div is cleared while generation is in progress. 3. When done, a download link reappears inside `#exportFile a`. 4. The CSV is fetched via `/review_porter/export//goodreads_export.csv`. Expect repeated 404s before the 200 arrives (Goodreads generates it async).