Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census
Abstract
The study audits AI venue recommendations across Bali markets, finding that visibility depends on documentation and ratings, with staleness rather than fabrication being the main failure mode.
AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.
Community
We pre-registered the design before collection: complete enumeration of Canggu and Ubud (4,776 venues), 2,208 traveler-style queries, 4 assistants, one week, 12,439 venue mentions resolved with a validated matching pipeline. Known limitations: one geography, one time window, and prompt phrasing effects (repeated identical queries overlap only 22 to 45%, which we report). Two of our hypotheses failed: star rating is null at the entry margin, and POI-database presence tested null at both margins. Happy to answer anything about the method.
The plain-language Report: https://norly.co/research/norly-invisible-to-ai-2026.pdf
Author here. Short version: we enumerated all 4,776 food/drink venues in two Bali markets, ran 2,208 persona-style queries against ChatGPT, Claude, Gemini and Perplexity (search-grounded APIs) over a week, resolved every mention back to the census, and modeled which venue-side factors predict appearing in answers vs being named first.
Findings I found most surprising: (1) 85.6% of venues never appear at all, 72.6% even among 50+ review venues; (2) retrieval and ranking are separable margins with different predictors, star rating is null for retrieval and only matters for rank; (3) fabrication is ~0.08% but stale (closed) venues get recommended 93 times, so the practical failure mode is freshness of the web the models read; (4) cross-engine top-20 agreement is 33-54%, and rerun stability is 22-45%, but a two-week retest shows no temporal drift.
Pre-registered, all 96 queries in the appendix, error rates for extraction/matching reported, and two null results against our own hypotheses. Funded by my company (review-management SaaS), disclosed in the paper. Glad to discuss method; the census-as-denominator design is the part I think generalizes.
Report: https://norly.co/research
Get this paper in your agent:
hf papers read 2608.07069 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper