Fix bot detection causing blank page from Goodreads

Goodreads returns an empty HTML page to headless browsers it identifies
as automated (navigator.webdriver=true, AutomationControlled feature).
This caused the Export Library button to never be found.

- Launch Chromium with --disable-blink-features=AutomationControlled
- Set a realistic user agent and viewport on the browser context
- Strip navigator.webdriver via an init script
- Guard the logged_in check so a blank page isn't mistaken for a session
- Wait for networkidle before querying the export button, and use
  wait_for(visible) instead of is_visible() to tolerate JS render delay
- Add CSS class fallback (button.js-LibraryExport) if role lookup fails
- Add AGENTS.md documenting the bot detection pattern and export flow

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Edward Betts 2026-02-18 11:53:42 +00:00
parent ae932b26b9
commit 56dff875e7
2 changed files with 88 additions and 18 deletions

48
AGENTS.md Normal file
View file

@ -0,0 +1,48 @@
# Agent Notes
## Bot detection (blank page symptom)
Goodreads returns `<html><head></head><body></body></html>` to headless browsers
it identifies as automated. Symptoms:
- `page.content()` → bare empty HTML skeleton
- `page.title()` → empty string
- Screenshot is fully white
- `.exportBooks` div count is 0
**Root cause:** Playwright's default headless Chromium exposes `navigator.webdriver = true`
and announces itself via the `AutomationControlled` Blink feature.
**Fix applied:**
- Launch with `--disable-blink-features=AutomationControlled`
- Set a realistic `user_agent` and `viewport` on the context
- `context.add_init_script(...)` to set `navigator.webdriver = undefined`
- Guard the `logged_in` check: a blank page must not be treated as "logged in"
## Goodreads Import/Export page (`/review/import`)
The "Export Library" button is present in the static HTML as:
```html
<button class='gr-form--compact__submitButton js-LibraryExport'
data-fileListId='exportFile'
data-statusId='exportStatusText'
data-userid='...'>Export Library</button>
```
### Why `get_by_role("button", name="Export Library")` can fail
`page.goto()` only waits for the `load` event. JavaScript may still be running
(React hydration, etc.) when the button is queried, so `count()` can return 0
or `is_visible()` can return `False` even though the button exists in the DOM.
**Fix applied:** call `page.wait_for_load_state("networkidle")` after `goto`,
then use `wait_for(state="visible", timeout=10s)` instead of `is_visible()`.
A CSS-class fallback (`button.js-LibraryExport`) is also tried in case the
accessible-name lookup fails.
### Export flow
1. Click "Export Library" → Goodreads POSTs to `/review_porter/export/<userid>`
2. The `#exportFile` div is cleared while generation is in progress.
3. When done, a download link reappears inside `#exportFile a`.
4. The CSV is fetched via `/review_porter/export/<userid>/goodreads_export.csv`.
Expect repeated 404s before the 200 arrives (Goodreads generates it async).