goodreads-backup/AGENTS.md
Edward Betts 56dff875e7 Fix bot detection causing blank page from Goodreads
Goodreads returns an empty HTML page to headless browsers it identifies
as automated (navigator.webdriver=true, AutomationControlled feature).
This caused the Export Library button to never be found.

- Launch Chromium with --disable-blink-features=AutomationControlled
- Set a realistic user agent and viewport on the browser context
- Strip navigator.webdriver via an init script
- Guard the logged_in check so a blank page isn't mistaken for a session
- Wait for networkidle before querying the export button, and use
  wait_for(visible) instead of is_visible() to tolerate JS render delay
- Add CSS class fallback (button.js-LibraryExport) if role lookup fails
- Add AGENTS.md documenting the bot detection pattern and export flow

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-18 11:53:42 +00:00

2 KiB

Agent Notes

Bot detection (blank page symptom)

Goodreads returns <html><head></head><body></body></html> to headless browsers it identifies as automated. Symptoms:

  • page.content() → bare empty HTML skeleton
  • page.title() → empty string
  • Screenshot is fully white
  • .exportBooks div count is 0

Root cause: Playwright's default headless Chromium exposes navigator.webdriver = true and announces itself via the AutomationControlled Blink feature.

Fix applied:

  • Launch with --disable-blink-features=AutomationControlled
  • Set a realistic user_agent and viewport on the context
  • context.add_init_script(...) to set navigator.webdriver = undefined
  • Guard the logged_in check: a blank page must not be treated as "logged in"

Goodreads Import/Export page (/review/import)

The "Export Library" button is present in the static HTML as:

<button class='gr-form--compact__submitButton js-LibraryExport'
        data-fileListId='exportFile'
        data-statusId='exportStatusText'
        data-userid='...'>Export Library</button>

Why get_by_role("button", name="Export Library") can fail

page.goto() only waits for the load event. JavaScript may still be running (React hydration, etc.) when the button is queried, so count() can return 0 or is_visible() can return False even though the button exists in the DOM.

Fix applied: call page.wait_for_load_state("networkidle") after goto, then use wait_for(state="visible", timeout=10s) instead of is_visible(). A CSS-class fallback (button.js-LibraryExport) is also tried in case the accessible-name lookup fails.

Export flow

  1. Click "Export Library" → Goodreads POSTs to /review_porter/export/<userid>
  2. The #exportFile div is cleared while generation is in progress.
  3. When done, a download link reappears inside #exportFile a.
  4. The CSV is fetched via /review_porter/export/<userid>/goodreads_export.csv. Expect repeated 404s before the 200 arrives (Goodreads generates it async).