← back to Kravet Sheet Sync 2026 04 20
__pycache__/kravet_image_catalog.cpython-314.pyc
89 lines
+
UK�j�% � � � R t ^ RIt^ RIt^ RIt^ RIt^ RIt^ RIt^ RIt^ RIt^ RI t ^ RI
t^ RIt^ RI
t^ RIHtHt RR/tRtRtRR.t]P* ! 4 tRsR tR
tRR ltRR ltR
tR tR tR tR t ]!R8X d
] ! 4 R# R# )u�
Build per-SKU image catalog for all Kravet wallcoverings.
For each SKU:
1) Query Algolia API → get canonical URL + primary/full_repeat Brandfolder URLs
2) Fetch product page HTML → extract additional sku-scoped brandfolder URLs
(room settings, detail shots, larger sizes)
3) Append one JSONL line per URL to kravet_image_catalog.jsonl
Features:
- Resumable: tracks completed SKUs in .done set
- Concurrent: 12 workers (tunable)
- Algolia API key is scraped from the search page; re-fetched if 401
- Rate-aware: retries with backoff on 429/5xx
After completion, rows are bulk-COPYd into dw_unified.kravet_sku_images.
N)�ThreadPoolExecutor�as_completedz
User-Agentz,Mozilla/5.0 (Macintosh; Intel Mac OS X 14_0)�
M9TBUM1WAE�$kravet_production_kravet_us_products�sshzroot@45.61.58.125c �� � Rp \ P P \ P P V \ R7 ^R7 P 4 P
RR4 p\ P ! RV\ P 4 pV'