Goodreads returns an empty HTML page to headless browsers it identifies as automated (navigator.webdriver=true, AutomationControlled feature). This caused the Export Library button to never be found. - Launch Chromium with --disable-blink-features=AutomationControlled - Set a realistic user agent and viewport on the browser context - Strip navigator.webdriver via an init script - Guard the logged_in check so a blank page isn't mistaken for a session - Wait for networkidle before querying the export button, and use wait_for(visible) instead of is_visible() to tolerate JS render delay - Add CSS class fallback (button.js-LibraryExport) if role lookup fails - Add AGENTS.md documenting the bot detection pattern and export flow Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
48 lines
2 KiB
Markdown
48 lines
2 KiB
Markdown
# Agent Notes
|
|
|
|
## Bot detection (blank page symptom)
|
|
|
|
Goodreads returns `<html><head></head><body></body></html>` to headless browsers
|
|
it identifies as automated. Symptoms:
|
|
- `page.content()` → bare empty HTML skeleton
|
|
- `page.title()` → empty string
|
|
- Screenshot is fully white
|
|
- `.exportBooks` div count is 0
|
|
|
|
**Root cause:** Playwright's default headless Chromium exposes `navigator.webdriver = true`
|
|
and announces itself via the `AutomationControlled` Blink feature.
|
|
|
|
**Fix applied:**
|
|
- Launch with `--disable-blink-features=AutomationControlled`
|
|
- Set a realistic `user_agent` and `viewport` on the context
|
|
- `context.add_init_script(...)` to set `navigator.webdriver = undefined`
|
|
- Guard the `logged_in` check: a blank page must not be treated as "logged in"
|
|
|
|
## Goodreads Import/Export page (`/review/import`)
|
|
|
|
The "Export Library" button is present in the static HTML as:
|
|
```html
|
|
<button class='gr-form--compact__submitButton js-LibraryExport'
|
|
data-fileListId='exportFile'
|
|
data-statusId='exportStatusText'
|
|
data-userid='...'>Export Library</button>
|
|
```
|
|
|
|
### Why `get_by_role("button", name="Export Library")` can fail
|
|
|
|
`page.goto()` only waits for the `load` event. JavaScript may still be running
|
|
(React hydration, etc.) when the button is queried, so `count()` can return 0
|
|
or `is_visible()` can return `False` even though the button exists in the DOM.
|
|
|
|
**Fix applied:** call `page.wait_for_load_state("networkidle")` after `goto`,
|
|
then use `wait_for(state="visible", timeout=10s)` instead of `is_visible()`.
|
|
A CSS-class fallback (`button.js-LibraryExport`) is also tried in case the
|
|
accessible-name lookup fails.
|
|
|
|
### Export flow
|
|
|
|
1. Click "Export Library" → Goodreads POSTs to `/review_porter/export/<userid>`
|
|
2. The `#exportFile` div is cleared while generation is in progress.
|
|
3. When done, a download link reappears inside `#exportFile a`.
|
|
4. The CSV is fetched via `/review_porter/export/<userid>/goodreads_export.csv`.
|
|
Expect repeated 404s before the 200 arrives (Goodreads generates it async).
|