Goodreads returns an empty HTML page to headless browsers it identifies as automated (navigator.webdriver=true, AutomationControlled feature). This caused the Export Library button to never be found. - Launch Chromium with --disable-blink-features=AutomationControlled - Set a realistic user agent and viewport on the browser context - Strip navigator.webdriver via an init script - Guard the logged_in check so a blank page isn't mistaken for a session - Wait for networkidle before querying the export button, and use wait_for(visible) instead of is_visible() to tolerate JS render delay - Add CSS class fallback (button.js-LibraryExport) if role lookup fails - Add AGENTS.md documenting the bot detection pattern and export flow Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2 KiB
Agent Notes
Bot detection (blank page symptom)
Goodreads returns <html><head></head><body></body></html> to headless browsers
it identifies as automated. Symptoms:
page.content()→ bare empty HTML skeletonpage.title()→ empty string- Screenshot is fully white
.exportBooksdiv count is 0
Root cause: Playwright's default headless Chromium exposes navigator.webdriver = true
and announces itself via the AutomationControlled Blink feature.
Fix applied:
- Launch with
--disable-blink-features=AutomationControlled - Set a realistic
user_agentandviewporton the context context.add_init_script(...)to setnavigator.webdriver = undefined- Guard the
logged_incheck: a blank page must not be treated as "logged in"
Goodreads Import/Export page (/review/import)
The "Export Library" button is present in the static HTML as:
<button class='gr-form--compact__submitButton js-LibraryExport'
data-fileListId='exportFile'
data-statusId='exportStatusText'
data-userid='...'>Export Library</button>
Why get_by_role("button", name="Export Library") can fail
page.goto() only waits for the load event. JavaScript may still be running
(React hydration, etc.) when the button is queried, so count() can return 0
or is_visible() can return False even though the button exists in the DOM.
Fix applied: call page.wait_for_load_state("networkidle") after goto,
then use wait_for(state="visible", timeout=10s) instead of is_visible().
A CSS-class fallback (button.js-LibraryExport) is also tried in case the
accessible-name lookup fails.
Export flow
- Click "Export Library" → Goodreads POSTs to
/review_porter/export/<userid> - The
#exportFilediv is cleared while generation is in progress. - When done, a download link reappears inside
#exportFile a. - The CSV is fetched via
/review_porter/export/<userid>/goodreads_export.csv. Expect repeated 404s before the 200 arrives (Goodreads generates it async).