diff --git a/README.md b/README.md index 409b722..9abd231 100644 --- a/README.md +++ b/README.md @@ -62,11 +62,11 @@ Playwright launches Chromium with `headless=False`. Complete sign-in and any CAP `--profile`, `--archive`, `--config`, and `--login-timeout` customize locations and login waiting time. Profile/archive paths are relative to the working directory. Do not run two browser processes with the same profile. -The command checks the delivered-order list through Ocado's explicit end-of-list marker, but filters out completed orders before loading their receipts. `--limit` counts new orders after this filtering. `--order-id` explicitly selects an old order for retry, refreshing, or importing previously excluded items; `--refresh` alone does not reselect completed orders. It reads each order's authenticated `/api/order/v6/orders/ID/decorated` JSON response as the browser loads it. It does not parse prices or quantities from rendered HTML or use the live trolley. Accepted substitutions are included; unavailable originals and rejected replacements are excluded. Discontinued products work without a product-page link. +The command checks the newest orders first and stops loading history once a batch contains a completed order. On the first run, it continues through Ocado's explicit end-of-list marker. Previously discovered incomplete orders remain eligible for retry, and completed orders are filtered out before loading receipts. Explicit `--order-id` selections continue searching until all requested orders are found or the list ends. `--limit` counts new orders after this filtering. `--order-id` explicitly selects an old order for retry, refreshing, or importing previously excluded items; `--refresh` alone does not reselect completed orders. It reads each order's authenticated `/api/order/v6/orders/ID/decorated` JSON response as the browser loads it. It does not parse prices or quantities from rendered HTML or use the live trolley. Accepted substitutions are included; unavailable originals and rejected replacements are excluded. Discontinued products work without a product-page link. Delivered item quantities and paid line prices must reconcile with the API order summary before importing. Delivery, bag charges and order-level credits are not stock items. Exact expiry dates and delivery dates come from the API, so historical orders do not depend on relative UI labels such as “Expired” or “Tuesday”. -Normalized receipts are saved to `history/ORDER_ID.json`, excluding account/payment details. Successful downloads are reused unless `--refresh` is specified. `history/orders.json` keeps the full manifest even when you use `--limit` or `--order-id`. Failed orders are listed in `history/download-errors.json` and retried on the next browser run. +Normalized receipts are saved to `history/ORDER_ID.json`, excluding account/payment details. Successful downloads are reused unless `--refresh` is specified. `history/orders.json` retains all previously discovered orders and merges in each new batch, including when you use `--limit` or `--order-id`. Failed orders are listed in `history/download-errors.json` and retried on the next browser run. ## Shelf-stable selection @@ -146,16 +146,16 @@ Tests use sanitized real API data and a fake Grocy API. They cover delivered ord .venv/bin/ocado-grocy-off # preview, no Grocy writes .venv/bin/ocado-grocy-off --apply # apply clear matches .venv/bin/ocado-grocy-off --offline # repeat a preview from cached data -.venv/bin/ocado-grocy-off --refresh-cache # download a fresh catalog +.venv/bin/ocado-grocy-off --refresh-cache # download fresh search responses ``` Set `[openfoodfacts].contact` in `config.toml` (or pass `--contact`) for the identifying User-Agent required by [Open Food Facts](https://openfoodfacts.github.io/openfoodfacts-server/api/). Requests are spaced at least 6.2 seconds apart. Run one enrichment process at a time. -The script reads the import journal for this Grocy server and matches live imported products against brand catalogs from the official [Search-a-licious API](https://openfoodfacts.github.io/search-a-licious/users/ref-openapi/). Large searches are split to avoid truncated results. Recognized household and personal-care categories are skipped. Missing data, different pack sizes, name/variant differences and multiple matching barcodes require review. Automatic matching requires the same brand, meaningful name words and metric pack size, including multipack structure, plus a valid GTIN check digit. Previously established Open Food Facts barcode matches are retained on reruns, including manually reviewed matches. Current product details are fetched again and checked before applying a catalog match. +The script reads the import journal for this Grocy server and searches by brand and product-name words for each eligible imported product using the official [Search-a-licious API](https://openfoodfacts.github.io/search-a-licious/users/ref-openapi/). Normal imports search only products from the orders being processed; queued failed enrichments are retried separately. Each product search reads at most 100 candidates, without paginating through whole brand catalogs. Truncated results require review and cannot produce automatic matches. `--limit` applies before searches, and non-food products do not trigger searches. Recognized household and personal-care categories are skipped. Missing data, different pack sizes, name/variant differences and multiple matching barcodes require review. Automatic matching requires the same brand, meaningful name words and metric pack size, including multipack structure, plus a valid GTIN check digit. Previously established Open Food Facts barcode matches are retained on reruns, including manually reviewed matches. Current product details are fetched again and checked before applying a search match. Writes add a barcode to Grocy's barcode table and an `openfoodfacts` section to the product description. Available nutrition, ingredient/allergen information, labels, packaging, serving size and image/source links are retained with attribution. Nutrition values keep their source units and basis; they are not copied into Grocy's calories-per-stock-unit field. Product titles, stock, purchase dates and stock notes are never edited. Deleted stock is never recreated. Product descriptions and barcodes are snapshotted before applying, and reruns are idempotent. Future Ocado metadata refreshes preserve the enrichment even after Grocy sanitizes the description HTML. -Private caches, snapshots and reports live in `history/openfoodfacts/`. `report.json` contains the preview and up to ten ranked candidates per product; `applied-report.json` records writes and failures. The catalog is only a source of candidates: Open Food Facts is crowdsourced, and matching by name cannot prove which physical barcode was on a previously purchased pack. Ambiguous matches are left unchanged. Check the current package for authoritative allergen information. +Private caches, snapshots and reports live in `history/openfoodfacts/`. `report.json` contains the preview and up to ten ranked candidates per product; `applied-report.json` records writes and failures. Search results are only candidates: Open Food Facts is crowdsourced, and matching by name cannot prove which physical barcode was on a previously purchased pack. Ambiguous matches are left unchanged. Check the current package for authoritative allergen information. For a manually verified match, create a JSON mapping of **Grocy product ID** to **barcode string** (retain leading zeroes), then run: diff --git a/ocado_grocy/history.py b/ocado_grocy/history.py index 45a267e..b01a260 100644 --- a/ocado_grocy/history.py +++ b/ocado_grocy/history.py @@ -66,7 +66,9 @@ def wait_for_orders(page, timeout, echo): raise ImportError("Timed out waiting for order history; rerun to reuse the saved browser session") -def discover_orders(page, echo): +def discover_orders(page, echo, *, completed=(), order_ids=()): + completed = set(completed) + requested = set(order_ids) last_count = -1 stable = 0 for _ in range(500): @@ -76,8 +78,17 @@ def discover_orders(page, echo): else: stable = 0 echo(f"Loaded {count} order links...") + links = order_links(page.content()) + visible = {link.order_id for link in links} + # Ocado lists newest orders first. Keep the whole visible batch so + # incomplete orders alongside the completion boundary remain eligible. + if requested: + if requested <= visible: + return links + elif completed & visible: + return links if page.locator('[data-test="order-list-no-more-orders-label"]').count(): - return order_links(page.content()) + return links if stable >= 10: raise ImportError("Order history stopped loading before Ocado's end-of-list marker; retry after checking the browser") last_count = count @@ -109,8 +120,8 @@ def download_orders(profile, archive, *, since=None, limit=None, login_only=Fals if login_only: echo("Signed-in browser session saved.") return [], [] - links = discover_orders(page, echo) - private_json(archive / "orders.json", [asdict(link) for link in links]) + links = discover_orders(page, echo, completed=completed, order_ids=order_ids) + links = save_order_manifest(archive, links) links = select_orders(links, since, limit, order_ids, completed) receipts, failures = [], [] for index, link in enumerate(links,1): @@ -161,6 +172,17 @@ def download_orders(profile, archive, *, since=None, limit=None, login_only=Fals context.close() +def save_order_manifest(archive, links): + """Retain older history and pending imports after an incremental scan.""" + path = archive / "orders.json" + saved = [OrderLink(**row) for row in json.loads(path.read_text())] if path.exists() else [] + merged = {link.order_id: link for link in saved} + merged.update((link.order_id, link) for link in links) + links = sorted(merged.values(), key=lambda link: (link.purchased_date, link.order_id), reverse=True) + private_json(path, [asdict(link) for link in links]) + return links + + def select_orders(links, since=None, limit=None, order_ids=(), completed=()): missing = set(order_ids) - {link.order_id for link in links} if missing: diff --git a/ocado_grocy/openfoodfacts.py b/ocado_grocy/openfoodfacts.py index d50a47f..629fe02 100644 --- a/ocado_grocy/openfoodfacts.py +++ b/ocado_grocy/openfoodfacts.py @@ -99,42 +99,31 @@ class OpenFoodFacts: private_json(target, data) return data - def catalog(self, brands, offline=False): - tags = sorted({brand_key(b).replace(' ', '-') for b in brands if b}) - if not tags: - return [] - - def group(tags): - query = 'brands_tags:(' + ' OR '.join(tags) + ')' - result = [] - page = 1 - while True: - data = self.fetch('https://search.openfoodfacts.org/search', {'q':query, 'page_size':1000, 'page':page, 'fields':FIELDS}, offline=offline) - if data.get('timed_out') or data.get('warnings'): - raise ValueError('Incomplete catalog search; retry later') - if not data.get('is_count_exact', False): - if len(tags) == 1: - raise ValueError('Brand catalog exceeds search limit: ' + tags[0]) - middle = len(tags) // 2 - return group(tags[:middle]) + group(tags[middle:]) - for hit in data.get('hits', []): - hit = dict(hit) - if isinstance(hit.get('brands'), list): - hit['brands'] = ','.join(hit['brands']) - result.append(hit) - click.echo(f"Catalog ({len(tags)} brands) page {page}: {len(result)} of {data.get('count')} products") - if page >= data.get('page_count', 1): - if len(result) != data.get('count'): - raise ValueError('Incomplete catalog response') - return result - page += 1 - - # OFF also retains apostrophes as tag separators (e.g. nairn-s). - alternatives = sorted({re.sub(r'[^a-z0-9]+', '-', unicodedata.normalize('NFKD', b).encode('ascii', 'ignore').decode().lower()).strip('-') for b in brands if b} - set(tags)) - result = group(tags) - if alternatives: - result += group(alternatives) - return list({str(p['code']):p for p in result}.values()) + def search(self, product, metadata, offline=False): + """Fetch one bounded page for this product, never a whole brand catalog.""" + brand = metadata.get('ocado', {}).get('brand', '') + terms = sorted(tokens(product['name'], brand)) + tags = {brand_key(brand).replace(' ', '-'), + re.sub(r'[^a-z0-9]+', '-', unicodedata.normalize('NFKD', brand) + .encode('ascii', 'ignore').decode().lower()).strip('-')} + if not brand or not terms: + return {}, {'query': '', 'search_count': 0, 'search_complete': True} + query = 'brands_tags:(' + ' OR '.join(sorted(tags)) + ') ' + ' '.join(terms) + data = self.fetch('https://search.openfoodfacts.org/search', + {'q': query, 'page_size': 100, 'page': 1, 'fields': FIELDS, 'langs': 'en'}, offline=offline) + if data.get('errors') or data.get('timed_out') or data.get('warnings'): + raise ValueError('Incomplete product search; retry later') + hits = data.get('hits', []) + complete = bool(data.get('is_count_exact')) and data.get('count') == len(hits) + candidates = {} + for hit in hits: + hit = dict(hit) + hit['code'] = str(hit['code']) + if isinstance(hit.get('brands'), list): + hit['brands'] = ','.join(hit['brands']) + candidates[hit['code']] = hit + return candidates, {'query': query, 'search_count': data.get('count'), + 'search_complete': complete} def product(self, code, offline=False): data = self.fetch(f'/api/v3.6/product/{code}.json', {'fields':FIELDS + ',nutrition'}, offline=offline) @@ -184,7 +173,7 @@ def apply_match(client, product, candidate, barcodes): @click.option('--apply', is_flag=True, help='Apply unambiguous matches and explicitly reviewed mappings.') @click.option('--offline', is_flag=True, help='Use cached Open Food Facts responses only.') @click.option('--mapping', type=click.Path(exists=True, path_type=Path), help='Reviewed JSON mapping of Grocy product ID to barcode.') -@click.option('--limit', type=int, help='Limit product comparisons after downloading the catalog.') +@click.option('--limit', type=int, help='Search at most this many eligible imported products.') @click.option('--order-id', 'order_ids', multiple=True, help='Limit to products imported from these orders.') def main(config, cache, contact, refresh_cache, apply, offline, mapping, limit, order_ids): """Match Ocado imports to Open Food Facts. Defaults to a read-only preview.""" @@ -209,17 +198,6 @@ def run_openfoodfacts(client, products, cache, contact, *, apply=True, offline=F mappings = mappings or {} if not isinstance(mappings, dict) or any(not isinstance(v, str) or not valid_barcode(v) for v in mappings.values()): raise click.UsageError('Mapping must be a JSON object with barcode strings and valid GTIN check digits') - def needs_search(product): - metadata = read_metadata(product.get('description')) - code = metadata.get('openfoodfacts', {}).get('code') - return not mappings.get(str(product['id'])) and not (valid_barcode(code) and code == metadata.get('Barcode')) - brands = {read_metadata(p.get('description')).get('ocado', {}).get('brand') for p in products if needs_search(p)} - catalog = off.catalog(brands, offline) - private_json(cache / 'catalog.json', catalog) - by_brand = {} - for candidate in catalog: - for brand in candidate.get('brands', '').split(','): - by_brand.setdefault(brand_key(brand), {})[str(candidate['code'])] = candidate report = {'created_at':datetime.now(timezone.utc).isoformat(), 'apply':apply, 'products':[]} if apply: private_json(cache / ('before-' + datetime.now().strftime('%Y%m%d-%H%M%S') + '.json'), {'products':products, 'barcodes':barcodes}) @@ -245,29 +223,30 @@ def run_openfoodfacts(client, products, cache, contact, *, apply=True, offline=F code = established method = 'existing_match' if code: - candidate = next((c for c in catalog if str(c.get('code')) == str(code)), None) or off.product(str(code), offline) + candidate = off.product(str(code), offline) row['match_method'] = method row['candidates'] = [candidate_info(product, metadata, candidate)] else: - query = 'brand catalog: ' + ocado.get('brand', '') - result = {'count':len(catalog)} - candidates = by_brand.get(brand_key(ocado.get('brand', '')), {}) - result['count'] = len(candidates) + click.echo(f"Searching Open Food Facts for {product['name']} ({row['size']})...") + candidates, search = off.search(product, metadata, offline) ranked, candidate = choose_match(product, metadata, candidates) - row.update(query=query, search_count=result.get('count'), candidates=ranked[:10]) + row.update(search, candidates=ranked[:10]) + # A capped result cannot establish that a barcode is unique. + if not search['search_complete']: + candidate = None row['match_method'] = 'exact_brand_name_pack' if candidate: - # Catalog is indexed separately; fetch current detail before writes. + # Search is indexed separately; fetch current detail before writes. if apply: fresh = off.product(str(candidate['code']), offline) row['current_candidate'] = candidate_info(product, metadata, fresh) if not code and not row['current_candidate']['exact']: - raise ValueError('Current product details no longer match catalog') + raise ValueError('Current product details no longer match search result') candidate = fresh row['barcode'] = candidate['code'] row['status'] = apply_match(client, product, candidate, barcodes) if apply else 'matched' else: - row['status'] = 'review' if row['candidates'] else 'no_match' + row['status'] = 'review' if row['candidates'] or not row.get('search_complete', True) else 'no_match' except (requests.RequestException, ValueError, RuntimeError) as exc: row.update(status='error', error=str(exc)) report['products'].append(row) diff --git a/tests/test_history.py b/tests/test_history.py index c523a60..03fd3bc 100644 --- a/tests/test_history.py +++ b/tests/test_history.py @@ -5,7 +5,8 @@ from pathlib import Path import pytest -from ocado_grocy.history import order_links,private_json,cached_orders,select_orders,OrderLink +from ocado_grocy.history import (order_links,private_json,cached_orders,select_orders,OrderLink, + discover_orders,save_order_manifest) from ocado_grocy.pantry import classify from ocado_grocy.receipt import ImportError,Item,parse_ocado_order from ocado_grocy.grocy import (Journal,import_receipt,description,refresh_imported_products, @@ -15,6 +16,76 @@ from test_grocy import FakeGrocy,receipt FIXTURE=Path(__file__).parent/'fixtures'/'history_receipt.json' +class HistoryPage: + def __init__(self, batches, end_marker=True): + self.batches = batches + self.index = 0 + self.end_marker = end_marker + self.mouse = self + + def locator(self, selector): + page = self + class Locator: + def count(self): + if selector == 'a[href*="/orders/"]': + return len(page.batches[page.index]) + return int(page.end_marker and page.index == len(page.batches) - 1) + @property + def last(self): + return self + def scroll_into_view_if_needed(self): + pass + return Locator() + + def get_by_role(self, *args, **kwargs): + return self.locator('button') + + def content(self): + return '' + ''.join( + f'Sep 8, 2026 Delivered' + for order in self.batches[self.index]) + '' + + def wheel(self, *args): + self.index = min(self.index + 1, len(self.batches) - 1) + + def wait_for_timeout(self, milliseconds): + pass + + +@pytest.mark.parametrize('batches,completed,expected_index', [ + ([['3', '2'], ['3', '2', '1']], {'2', '1'}, 0), + ([['3'], ['3', '2'], ['3', '2', '1']], {'2', '1'}, 1), + ([['3'], ['3', '2'], ['3', '2', '1']], set(), 2), +]) +def test_discovery_stops_at_completed_batch_or_end(batches, completed, expected_index): + page = HistoryPage(batches) + links = discover_orders(page, lambda message: None, completed=completed) + assert page.index == expected_index + assert [link.order_id for link in links] == batches[expected_index] + + +def test_explicit_discovery_searches_past_completed_orders(): + page = HistoryPage([['3'], ['3', '2'], ['3', '2', '1']], end_marker=False) + links = discover_orders(page, lambda message: None, completed={'3'}, order_ids=('2',)) + assert page.index == 1 + assert [link.order_id for link in links] == ['3', '2'] + + +def test_discovery_still_rejects_stalled_history(): + with pytest.raises(ImportError, match='stopped loading'): + discover_orders(HistoryPage([['3']], end_marker=False), lambda message: None) + + +def test_incremental_manifest_preserves_older_pending_orders(tmp_path): + old = [OrderLink('2', '2026-09-07', 'old-url'), OrderLink('1', '2026-09-06', 'url')] + save_order_manifest(tmp_path, old) + new = [OrderLink('3', '2026-09-08', 'url'), replace(old[0], url='updated-url')] + merged = save_order_manifest(tmp_path, new) + assert merged == new + old[1:] + assert [link.order_id for link in select_orders(merged, completed={'2'})] == ['3', '1'] + assert len(json.loads((tmp_path/'orders.json').read_text())) == 3 + + def test_storage_and_categories_select_without_a_product_list(): order=parse_ocado_order(json.loads(FIXTURE.read_text())) assert sum(classify(i).include for i in order.items)==9 diff --git a/tests/test_openfoodfacts.py b/tests/test_openfoodfacts.py index 6e3c14c..1e357db 100644 --- a/tests/test_openfoodfacts.py +++ b/tests/test_openfoodfacts.py @@ -83,19 +83,54 @@ def test_two_exact_barcodes_are_ambiguous(): assert match is None -def test_truncated_catalog_is_split_and_pages_must_be_complete(tmp_path): +def test_product_search_is_bounded_and_includes_name_and_brand(tmp_path): from ocado_grocy.openfoodfacts import OpenFoodFacts + calls = [] class OFF(OpenFoodFacts): def fetch(self, path, params, **kwargs): - if params['q'] == 'brands_tags:(a OR b)': - return {'is_count_exact':False, 'hits':[]} - code = '00000048' if params['q'] == 'brands_tags:(a)' else '00000055' - return {'is_count_exact':True, 'count':1, 'page_count':1, 'hits':[{'code':code,'brands':['a']}]} - off = OFF(tmp_path, 'test@example.com') - assert len(off.catalog({'a','b'})) == 2 - off.fetch = lambda *args, **kwargs: {'is_count_exact':True, 'count':2, 'page_count':1, 'hits':[]} - with pytest.raises(ValueError, match='Incomplete'): - off.catalog({'a'}) + calls.append(params) + return {'is_count_exact': True, 'count': 9000, + 'hits': [{**candidate(), 'brands': ['Marks & Spencer']}]} + candidates, result = OFF(tmp_path, 'test@example.com').search(product(), metadata()) + assert len(calls) == 1 + assert calls[0]['page_size'] == 100 and calls[0]['page'] == 1 + assert 'ginger' in calls[0]['q'] and 'snaps' in calls[0]['q'] + assert 'marks-spencer' in calls[0]['q'] + assert candidates[candidate()['code']]['brands'] == 'Marks & Spencer' + assert not result['search_complete'] + + +@pytest.mark.parametrize('complete', [True, False]) +def test_matching_searches_only_selected_products_and_requires_complete_results(tmp_path, monkeypatch, complete): + from ocado_grocy import openfoodfacts as module + searched = [] + class OFF: + def __init__(self, *args): pass + def search(self, p, m, offline): + searched.append(p['id']) + return {candidate()['code']: candidate()}, {'search_complete': complete} + class Client: + def get(self, path): + assert path == '/objects/product_barcodes' + return [] + monkeypatch.setattr(module, 'OpenFoodFacts', OFF) + nonfood = {**product(), 'id': 43} + nonfood['description'] = replace_metadata(nonfood['description'], + {**metadata(), 'ocado': {'brand': 'Miniml'}}) + result = module.run_openfoodfacts(Client(), [nonfood, product(), {**product(), 'id': 44}], + tmp_path, 'test@example.com', apply=False, limit=1) + assert searched == [42] + assert [r['status'] for r in result['products']] == [ + 'non_food', 'matched' if complete else 'review', 'not_searched'] + + +@pytest.mark.parametrize('error', [{'timed_out': True}, {'warnings': ['partial']}, {'errors': ['bad query']}]) +def test_failed_search_is_not_reported_as_no_match(tmp_path, error): + from ocado_grocy.openfoodfacts import OpenFoodFacts + off = OpenFoodFacts(tmp_path, 'test@example.com') + off.fetch = lambda *args, **kwargs: error + with pytest.raises(ValueError, match='Incomplete product search'): + off.search(product(), metadata()) def test_ms_food_brand_and_missing_barcode_recovery(): @@ -116,7 +151,7 @@ def test_established_barcode_match_survives_a_missing_catalog_entry(tmp_path, mo p['description'] = replace_metadata(p['description'], m) class OFF: def __init__(self, *a): pass - def catalog(self, *a): return [] + def search(self, *a): pytest.fail("Established match must not trigger search") def product(self, code, offline): assert code == candidate()['code'] return candidate()