Limit order discovery and Open Food Facts searches

This commit is contained in:
Edward Betts 2026-10-07 11:16:50 +01:00
parent 628e5c3823
commit 41772b4d1a
5 changed files with 185 additions and 78 deletions

View file

@ -62,11 +62,11 @@ Playwright launches Chromium with `headless=False`. Complete sign-in and any CAP
`--profile`, `--archive`, `--config`, and `--login-timeout` customize locations and login waiting time. Profile/archive paths are relative to the working directory. Do not run two browser processes with the same profile.
The command checks the delivered-order list through Ocado's explicit end-of-list marker, but filters out completed orders before loading their receipts. `--limit` counts new orders after this filtering. `--order-id` explicitly selects an old order for retry, refreshing, or importing previously excluded items; `--refresh` alone does not reselect completed orders. It reads each order's authenticated `/api/order/v6/orders/ID/decorated` JSON response as the browser loads it. It does not parse prices or quantities from rendered HTML or use the live trolley. Accepted substitutions are included; unavailable originals and rejected replacements are excluded. Discontinued products work without a product-page link.
The command checks the newest orders first and stops loading history once a batch contains a completed order. On the first run, it continues through Ocado's explicit end-of-list marker. Previously discovered incomplete orders remain eligible for retry, and completed orders are filtered out before loading receipts. Explicit `--order-id` selections continue searching until all requested orders are found or the list ends. `--limit` counts new orders after this filtering. `--order-id` explicitly selects an old order for retry, refreshing, or importing previously excluded items; `--refresh` alone does not reselect completed orders. It reads each order's authenticated `/api/order/v6/orders/ID/decorated` JSON response as the browser loads it. It does not parse prices or quantities from rendered HTML or use the live trolley. Accepted substitutions are included; unavailable originals and rejected replacements are excluded. Discontinued products work without a product-page link.
Delivered item quantities and paid line prices must reconcile with the API order summary before importing. Delivery, bag charges and order-level credits are not stock items. Exact expiry dates and delivery dates come from the API, so historical orders do not depend on relative UI labels such as “Expired” or “Tuesday”.
Normalized receipts are saved to `history/ORDER_ID.json`, excluding account/payment details. Successful downloads are reused unless `--refresh` is specified. `history/orders.json` keeps the full manifest even when you use `--limit` or `--order-id`. Failed orders are listed in `history/download-errors.json` and retried on the next browser run.
Normalized receipts are saved to `history/ORDER_ID.json`, excluding account/payment details. Successful downloads are reused unless `--refresh` is specified. `history/orders.json` retains all previously discovered orders and merges in each new batch, including when you use `--limit` or `--order-id`. Failed orders are listed in `history/download-errors.json` and retried on the next browser run.
## Shelf-stable selection
@ -146,16 +146,16 @@ Tests use sanitized real API data and a fake Grocy API. They cover delivered ord
.venv/bin/ocado-grocy-off # preview, no Grocy writes
.venv/bin/ocado-grocy-off --apply # apply clear matches
.venv/bin/ocado-grocy-off --offline # repeat a preview from cached data
.venv/bin/ocado-grocy-off --refresh-cache # download a fresh catalog
.venv/bin/ocado-grocy-off --refresh-cache # download fresh search responses
```
Set `[openfoodfacts].contact` in `config.toml` (or pass `--contact`) for the identifying User-Agent required by [Open Food Facts](https://openfoodfacts.github.io/openfoodfacts-server/api/). Requests are spaced at least 6.2 seconds apart. Run one enrichment process at a time.
The script reads the import journal for this Grocy server and matches live imported products against brand catalogs from the official [Search-a-licious API](https://openfoodfacts.github.io/search-a-licious/users/ref-openapi/). Large searches are split to avoid truncated results. Recognized household and personal-care categories are skipped. Missing data, different pack sizes, name/variant differences and multiple matching barcodes require review. Automatic matching requires the same brand, meaningful name words and metric pack size, including multipack structure, plus a valid GTIN check digit. Previously established Open Food Facts barcode matches are retained on reruns, including manually reviewed matches. Current product details are fetched again and checked before applying a catalog match.
The script reads the import journal for this Grocy server and searches by brand and product-name words for each eligible imported product using the official [Search-a-licious API](https://openfoodfacts.github.io/search-a-licious/users/ref-openapi/). Normal imports search only products from the orders being processed; queued failed enrichments are retried separately. Each product search reads at most 100 candidates, without paginating through whole brand catalogs. Truncated results require review and cannot produce automatic matches. `--limit` applies before searches, and non-food products do not trigger searches. Recognized household and personal-care categories are skipped. Missing data, different pack sizes, name/variant differences and multiple matching barcodes require review. Automatic matching requires the same brand, meaningful name words and metric pack size, including multipack structure, plus a valid GTIN check digit. Previously established Open Food Facts barcode matches are retained on reruns, including manually reviewed matches. Current product details are fetched again and checked before applying a search match.
Writes add a barcode to Grocy's barcode table and an `openfoodfacts` section to the product description. Available nutrition, ingredient/allergen information, labels, packaging, serving size and image/source links are retained with attribution. Nutrition values keep their source units and basis; they are not copied into Grocy's calories-per-stock-unit field. Product titles, stock, purchase dates and stock notes are never edited. Deleted stock is never recreated. Product descriptions and barcodes are snapshotted before applying, and reruns are idempotent. Future Ocado metadata refreshes preserve the enrichment even after Grocy sanitizes the description HTML.
Private caches, snapshots and reports live in `history/openfoodfacts/`. `report.json` contains the preview and up to ten ranked candidates per product; `applied-report.json` records writes and failures. The catalog is only a source of candidates: Open Food Facts is crowdsourced, and matching by name cannot prove which physical barcode was on a previously purchased pack. Ambiguous matches are left unchanged. Check the current package for authoritative allergen information.
Private caches, snapshots and reports live in `history/openfoodfacts/`. `report.json` contains the preview and up to ten ranked candidates per product; `applied-report.json` records writes and failures. Search results are only candidates: Open Food Facts is crowdsourced, and matching by name cannot prove which physical barcode was on a previously purchased pack. Ambiguous matches are left unchanged. Check the current package for authoritative allergen information.
For a manually verified match, create a JSON mapping of **Grocy product ID** to **barcode string** (retain leading zeroes), then run:

View file

@ -66,7 +66,9 @@ def wait_for_orders(page, timeout, echo):
raise ImportError("Timed out waiting for order history; rerun to reuse the saved browser session")
def discover_orders(page, echo):
def discover_orders(page, echo, *, completed=(), order_ids=()):
completed = set(completed)
requested = set(order_ids)
last_count = -1
stable = 0
for _ in range(500):
@ -76,8 +78,17 @@ def discover_orders(page, echo):
else:
stable = 0
echo(f"Loaded {count} order links...")
links = order_links(page.content())
visible = {link.order_id for link in links}
# Ocado lists newest orders first. Keep the whole visible batch so
# incomplete orders alongside the completion boundary remain eligible.
if requested:
if requested <= visible:
return links
elif completed & visible:
return links
if page.locator('[data-test="order-list-no-more-orders-label"]').count():
return order_links(page.content())
return links
if stable >= 10:
raise ImportError("Order history stopped loading before Ocado's end-of-list marker; retry after checking the browser")
last_count = count
@ -109,8 +120,8 @@ def download_orders(profile, archive, *, since=None, limit=None, login_only=Fals
if login_only:
echo("Signed-in browser session saved.")
return [], []
links = discover_orders(page, echo)
private_json(archive / "orders.json", [asdict(link) for link in links])
links = discover_orders(page, echo, completed=completed, order_ids=order_ids)
links = save_order_manifest(archive, links)
links = select_orders(links, since, limit, order_ids, completed)
receipts, failures = [], []
for index, link in enumerate(links,1):
@ -161,6 +172,17 @@ def download_orders(profile, archive, *, since=None, limit=None, login_only=Fals
context.close()
def save_order_manifest(archive, links):
"""Retain older history and pending imports after an incremental scan."""
path = archive / "orders.json"
saved = [OrderLink(**row) for row in json.loads(path.read_text())] if path.exists() else []
merged = {link.order_id: link for link in saved}
merged.update((link.order_id, link) for link in links)
links = sorted(merged.values(), key=lambda link: (link.purchased_date, link.order_id), reverse=True)
private_json(path, [asdict(link) for link in links])
return links
def select_orders(links, since=None, limit=None, order_ids=(), completed=()):
missing = set(order_ids) - {link.order_id for link in links}
if missing:

View file

@ -99,42 +99,31 @@ class OpenFoodFacts:
private_json(target, data)
return data
def catalog(self, brands, offline=False):
tags = sorted({brand_key(b).replace(' ', '-') for b in brands if b})
if not tags:
return []
def group(tags):
query = 'brands_tags:(' + ' OR '.join(tags) + ')'
result = []
page = 1
while True:
data = self.fetch('https://search.openfoodfacts.org/search', {'q':query, 'page_size':1000, 'page':page, 'fields':FIELDS}, offline=offline)
if data.get('timed_out') or data.get('warnings'):
raise ValueError('Incomplete catalog search; retry later')
if not data.get('is_count_exact', False):
if len(tags) == 1:
raise ValueError('Brand catalog exceeds search limit: ' + tags[0])
middle = len(tags) // 2
return group(tags[:middle]) + group(tags[middle:])
for hit in data.get('hits', []):
def search(self, product, metadata, offline=False):
"""Fetch one bounded page for this product, never a whole brand catalog."""
brand = metadata.get('ocado', {}).get('brand', '')
terms = sorted(tokens(product['name'], brand))
tags = {brand_key(brand).replace(' ', '-'),
re.sub(r'[^a-z0-9]+', '-', unicodedata.normalize('NFKD', brand)
.encode('ascii', 'ignore').decode().lower()).strip('-')}
if not brand or not terms:
return {}, {'query': '', 'search_count': 0, 'search_complete': True}
query = 'brands_tags:(' + ' OR '.join(sorted(tags)) + ') ' + ' '.join(terms)
data = self.fetch('https://search.openfoodfacts.org/search',
{'q': query, 'page_size': 100, 'page': 1, 'fields': FIELDS, 'langs': 'en'}, offline=offline)
if data.get('errors') or data.get('timed_out') or data.get('warnings'):
raise ValueError('Incomplete product search; retry later')
hits = data.get('hits', [])
complete = bool(data.get('is_count_exact')) and data.get('count') == len(hits)
candidates = {}
for hit in hits:
hit = dict(hit)
hit['code'] = str(hit['code'])
if isinstance(hit.get('brands'), list):
hit['brands'] = ','.join(hit['brands'])
result.append(hit)
click.echo(f"Catalog ({len(tags)} brands) page {page}: {len(result)} of {data.get('count')} products")
if page >= data.get('page_count', 1):
if len(result) != data.get('count'):
raise ValueError('Incomplete catalog response')
return result
page += 1
# OFF also retains apostrophes as tag separators (e.g. nairn-s).
alternatives = sorted({re.sub(r'[^a-z0-9]+', '-', unicodedata.normalize('NFKD', b).encode('ascii', 'ignore').decode().lower()).strip('-') for b in brands if b} - set(tags))
result = group(tags)
if alternatives:
result += group(alternatives)
return list({str(p['code']):p for p in result}.values())
candidates[hit['code']] = hit
return candidates, {'query': query, 'search_count': data.get('count'),
'search_complete': complete}
def product(self, code, offline=False):
data = self.fetch(f'/api/v3.6/product/{code}.json', {'fields':FIELDS + ',nutrition'}, offline=offline)
@ -184,7 +173,7 @@ def apply_match(client, product, candidate, barcodes):
@click.option('--apply', is_flag=True, help='Apply unambiguous matches and explicitly reviewed mappings.')
@click.option('--offline', is_flag=True, help='Use cached Open Food Facts responses only.')
@click.option('--mapping', type=click.Path(exists=True, path_type=Path), help='Reviewed JSON mapping of Grocy product ID to barcode.')
@click.option('--limit', type=int, help='Limit product comparisons after downloading the catalog.')
@click.option('--limit', type=int, help='Search at most this many eligible imported products.')
@click.option('--order-id', 'order_ids', multiple=True, help='Limit to products imported from these orders.')
def main(config, cache, contact, refresh_cache, apply, offline, mapping, limit, order_ids):
"""Match Ocado imports to Open Food Facts. Defaults to a read-only preview."""
@ -209,17 +198,6 @@ def run_openfoodfacts(client, products, cache, contact, *, apply=True, offline=F
mappings = mappings or {}
if not isinstance(mappings, dict) or any(not isinstance(v, str) or not valid_barcode(v) for v in mappings.values()):
raise click.UsageError('Mapping must be a JSON object with barcode strings and valid GTIN check digits')
def needs_search(product):
metadata = read_metadata(product.get('description'))
code = metadata.get('openfoodfacts', {}).get('code')
return not mappings.get(str(product['id'])) and not (valid_barcode(code) and code == metadata.get('Barcode'))
brands = {read_metadata(p.get('description')).get('ocado', {}).get('brand') for p in products if needs_search(p)}
catalog = off.catalog(brands, offline)
private_json(cache / 'catalog.json', catalog)
by_brand = {}
for candidate in catalog:
for brand in candidate.get('brands', '').split(','):
by_brand.setdefault(brand_key(brand), {})[str(candidate['code'])] = candidate
report = {'created_at':datetime.now(timezone.utc).isoformat(), 'apply':apply, 'products':[]}
if apply:
private_json(cache / ('before-' + datetime.now().strftime('%Y%m%d-%H%M%S') + '.json'), {'products':products, 'barcodes':barcodes})
@ -245,29 +223,30 @@ def run_openfoodfacts(client, products, cache, contact, *, apply=True, offline=F
code = established
method = 'existing_match'
if code:
candidate = next((c for c in catalog if str(c.get('code')) == str(code)), None) or off.product(str(code), offline)
candidate = off.product(str(code), offline)
row['match_method'] = method
row['candidates'] = [candidate_info(product, metadata, candidate)]
else:
query = 'brand catalog: ' + ocado.get('brand', '')
result = {'count':len(catalog)}
candidates = by_brand.get(brand_key(ocado.get('brand', '')), {})
result['count'] = len(candidates)
click.echo(f"Searching Open Food Facts for {product['name']} ({row['size']})...")
candidates, search = off.search(product, metadata, offline)
ranked, candidate = choose_match(product, metadata, candidates)
row.update(query=query, search_count=result.get('count'), candidates=ranked[:10])
row.update(search, candidates=ranked[:10])
# A capped result cannot establish that a barcode is unique.
if not search['search_complete']:
candidate = None
row['match_method'] = 'exact_brand_name_pack'
if candidate:
# Catalog is indexed separately; fetch current detail before writes.
# Search is indexed separately; fetch current detail before writes.
if apply:
fresh = off.product(str(candidate['code']), offline)
row['current_candidate'] = candidate_info(product, metadata, fresh)
if not code and not row['current_candidate']['exact']:
raise ValueError('Current product details no longer match catalog')
raise ValueError('Current product details no longer match search result')
candidate = fresh
row['barcode'] = candidate['code']
row['status'] = apply_match(client, product, candidate, barcodes) if apply else 'matched'
else:
row['status'] = 'review' if row['candidates'] else 'no_match'
row['status'] = 'review' if row['candidates'] or not row.get('search_complete', True) else 'no_match'
except (requests.RequestException, ValueError, RuntimeError) as exc:
row.update(status='error', error=str(exc))
report['products'].append(row)

View file

@ -5,7 +5,8 @@ from pathlib import Path
import pytest
from ocado_grocy.history import order_links,private_json,cached_orders,select_orders,OrderLink
from ocado_grocy.history import (order_links,private_json,cached_orders,select_orders,OrderLink,
discover_orders,save_order_manifest)
from ocado_grocy.pantry import classify
from ocado_grocy.receipt import ImportError,Item,parse_ocado_order
from ocado_grocy.grocy import (Journal,import_receipt,description,refresh_imported_products,
@ -15,6 +16,76 @@ from test_grocy import FakeGrocy,receipt
FIXTURE=Path(__file__).parent/'fixtures'/'history_receipt.json'
class HistoryPage:
def __init__(self, batches, end_marker=True):
self.batches = batches
self.index = 0
self.end_marker = end_marker
self.mouse = self
def locator(self, selector):
page = self
class Locator:
def count(self):
if selector == 'a[href*="/orders/"]':
return len(page.batches[page.index])
return int(page.end_marker and page.index == len(page.batches) - 1)
@property
def last(self):
return self
def scroll_into_view_if_needed(self):
pass
return Locator()
def get_by_role(self, *args, **kwargs):
return self.locator('button')
def content(self):
return '<html>' + ''.join(
f'<a href="/orders/{order}/details">Sep 8, 2026 Delivered</a>'
for order in self.batches[self.index]) + '</html>'
def wheel(self, *args):
self.index = min(self.index + 1, len(self.batches) - 1)
def wait_for_timeout(self, milliseconds):
pass
@pytest.mark.parametrize('batches,completed,expected_index', [
([['3', '2'], ['3', '2', '1']], {'2', '1'}, 0),
([['3'], ['3', '2'], ['3', '2', '1']], {'2', '1'}, 1),
([['3'], ['3', '2'], ['3', '2', '1']], set(), 2),
])
def test_discovery_stops_at_completed_batch_or_end(batches, completed, expected_index):
page = HistoryPage(batches)
links = discover_orders(page, lambda message: None, completed=completed)
assert page.index == expected_index
assert [link.order_id for link in links] == batches[expected_index]
def test_explicit_discovery_searches_past_completed_orders():
page = HistoryPage([['3'], ['3', '2'], ['3', '2', '1']], end_marker=False)
links = discover_orders(page, lambda message: None, completed={'3'}, order_ids=('2',))
assert page.index == 1
assert [link.order_id for link in links] == ['3', '2']
def test_discovery_still_rejects_stalled_history():
with pytest.raises(ImportError, match='stopped loading'):
discover_orders(HistoryPage([['3']], end_marker=False), lambda message: None)
def test_incremental_manifest_preserves_older_pending_orders(tmp_path):
old = [OrderLink('2', '2026-09-07', 'old-url'), OrderLink('1', '2026-09-06', 'url')]
save_order_manifest(tmp_path, old)
new = [OrderLink('3', '2026-09-08', 'url'), replace(old[0], url='updated-url')]
merged = save_order_manifest(tmp_path, new)
assert merged == new + old[1:]
assert [link.order_id for link in select_orders(merged, completed={'2'})] == ['3', '1']
assert len(json.loads((tmp_path/'orders.json').read_text())) == 3
def test_storage_and_categories_select_without_a_product_list():
order=parse_ocado_order(json.loads(FIXTURE.read_text()))
assert sum(classify(i).include for i in order.items)==9

View file

@ -83,19 +83,54 @@ def test_two_exact_barcodes_are_ambiguous():
assert match is None
def test_truncated_catalog_is_split_and_pages_must_be_complete(tmp_path):
def test_product_search_is_bounded_and_includes_name_and_brand(tmp_path):
from ocado_grocy.openfoodfacts import OpenFoodFacts
calls = []
class OFF(OpenFoodFacts):
def fetch(self, path, params, **kwargs):
if params['q'] == 'brands_tags:(a OR b)':
return {'is_count_exact':False, 'hits':[]}
code = '00000048' if params['q'] == 'brands_tags:(a)' else '00000055'
return {'is_count_exact':True, 'count':1, 'page_count':1, 'hits':[{'code':code,'brands':['a']}]}
off = OFF(tmp_path, 'test@example.com')
assert len(off.catalog({'a','b'})) == 2
off.fetch = lambda *args, **kwargs: {'is_count_exact':True, 'count':2, 'page_count':1, 'hits':[]}
with pytest.raises(ValueError, match='Incomplete'):
off.catalog({'a'})
calls.append(params)
return {'is_count_exact': True, 'count': 9000,
'hits': [{**candidate(), 'brands': ['Marks & Spencer']}]}
candidates, result = OFF(tmp_path, 'test@example.com').search(product(), metadata())
assert len(calls) == 1
assert calls[0]['page_size'] == 100 and calls[0]['page'] == 1
assert 'ginger' in calls[0]['q'] and 'snaps' in calls[0]['q']
assert 'marks-spencer' in calls[0]['q']
assert candidates[candidate()['code']]['brands'] == 'Marks & Spencer'
assert not result['search_complete']
@pytest.mark.parametrize('complete', [True, False])
def test_matching_searches_only_selected_products_and_requires_complete_results(tmp_path, monkeypatch, complete):
from ocado_grocy import openfoodfacts as module
searched = []
class OFF:
def __init__(self, *args): pass
def search(self, p, m, offline):
searched.append(p['id'])
return {candidate()['code']: candidate()}, {'search_complete': complete}
class Client:
def get(self, path):
assert path == '/objects/product_barcodes'
return []
monkeypatch.setattr(module, 'OpenFoodFacts', OFF)
nonfood = {**product(), 'id': 43}
nonfood['description'] = replace_metadata(nonfood['description'],
{**metadata(), 'ocado': {'brand': 'Miniml'}})
result = module.run_openfoodfacts(Client(), [nonfood, product(), {**product(), 'id': 44}],
tmp_path, 'test@example.com', apply=False, limit=1)
assert searched == [42]
assert [r['status'] for r in result['products']] == [
'non_food', 'matched' if complete else 'review', 'not_searched']
@pytest.mark.parametrize('error', [{'timed_out': True}, {'warnings': ['partial']}, {'errors': ['bad query']}])
def test_failed_search_is_not_reported_as_no_match(tmp_path, error):
from ocado_grocy.openfoodfacts import OpenFoodFacts
off = OpenFoodFacts(tmp_path, 'test@example.com')
off.fetch = lambda *args, **kwargs: error
with pytest.raises(ValueError, match='Incomplete product search'):
off.search(product(), metadata())
def test_ms_food_brand_and_missing_barcode_recovery():
@ -116,7 +151,7 @@ def test_established_barcode_match_survives_a_missing_catalog_entry(tmp_path, mo
p['description'] = replace_metadata(p['description'], m)
class OFF:
def __init__(self, *a): pass
def catalog(self, *a): return []
def search(self, *a): pytest.fail("Established match must not trigger search")
def product(self, code, offline):
assert code == candidate()['code']
return candidate()