Call · 15 min
StudyMattia Esposito26 September 202612 min read

Producer websites are open to everyone. One in six offers a spec sheet.

We examined, from the outside, the 249 websites listed by producers and exporters in ICE's directory for foreign buyers. The sites are open to everyone, search engines and AI assistants included. The spec sheet a buyer wants to see is offered by just one food producer in six.

In brief

19 of the 121 food websites (15.7%) offer a downloadable spec sheet between the home page and a product page, and another 7 show it on the page only. Among wineries, 13 out of 25 offer one; among other food producers, 6 out of 96. Every sheet was opened and read, one by one.

All 121 food websites let in Google, Bing and the search crawlers used by ChatGPT, Claude and Perplexity. The gap is on the page itself: in 18.2% the title is just the company name, 35.5% have no meta description, 43.8% don't use structured data to say who the company is, and 11.6% have an email address that AI assistants can't see.

26 of the 249 addresses in the directory (10.4%) don't lead to a working website: domains that no longer exist, error pages, sites under maintenance, a WordPress install that was never set up, and a domain now taken over by an online casino.

The study is by Itria AI (itria.io) and is independent of ICE: the directory is simply the public source the website addresses came from. The measurements were taken on 26 September 2026. The method, checks, corrections and the data for every site, without company names, are on this page and in the file you can download at the bottom, under a CC BY 4.0 licence. If you export, the practical pages start with the hands-on export guide for small food producers.

Where export deals really happen

In person, at trade fairs such as Anuga and SIAL, and through samples. The CBI, the Dutch government centre that helps exporters from middle- and low-income countries sell into Europe, says so in its guide on how to find European buyers of processed fruit and vegetables: most lasting partnerships still come from personal introductions, sample testing and factory visits.

“Processed food is a face-to-face business, but a website gives you global visibility” (CBI, guide to European buyers of processed fruit and vegetables, updated 19 September 2023)

Before the meeting comes the search. European buyers look for new suppliers on Google, typing the product name and words like “producer” or “supplier”, sometimes with the country of origin, and according to the CBI appearing on the first page for those searches is “very important”. When it's the producer who makes the first approach, the CBI recommends a phone call followed by an email.

On the website itself, the CBI says, product documentation should be available to buyers: full specifications, including size, weight, packaging, shelf life and nutritional values, which is the information a buyer wants to see and that helps them decide. As an example it points to a producer that publishes a PDF sheet for every product: that's the spec sheet this study went looking for.

The study measures what can be seen from the outside: what a buyer, a search engine and an AI assistant find on the website. How many export enquiries actually come in through the website is known only to the company receiving them. That figure is missing from the public sources we read, and we don't estimate it here.

What we measured: the buyer, the search engine, the AI assistant

A company website now has three readers: the foreign buyer who opens it, the search engine that indexes it, and the AI assistant that reads it on someone else's behalf. We measured five things the first one needs, and six that decide whether the other two can find the company and understand what it does.

MeasureWhat counts as a yesWhy it matters
English version

English hreflang, a link to /en/, an EN language switcher, or a home page already in English.

It's the foreign buyer's language.

Contact channel

A form with an email or message field, a mailto link, or a written address.

Without one, the buyer doesn't write.

Email readable by an AI assistant

The address in plain text in the served HTML, not obfuscated.

Assistants that read pages on the fly don't run JavaScript (Vercel).

Spec sheet

Downloadable: a product sheet as a PDF, image or web page, linked somewhere between the home page and a product page, and opened. On the page: a section headed technical sheet or technical data, with the fields below. A general catalogue doesn't count.

The CBI recommends keeping it on the website, available to buyers.

Terms for foreign buyers

An export or wholesale price list, a minimum order for resellers, an Incoterm, an ex-works price. Checked case by case.

Without them, every offer starts with an email asking a question.

robots.txt open to Google and Bing

Googlebot and Bingbot can read the home page.

Block them and you drop out of search results.

robots.txt open to AI search

OAI-SearchBot, PerplexityBot, Claude-SearchBot and Claude-User can read the home page.

OpenAI: sites that block OAI-SearchBot don't appear in ChatGPT's search answers.

Indexable home page

No noindex in the meta robots tag or the X-Robots-Tag header.

Google drops a page with noindex from its results altogether.

A title that says what the company does

The title contains something besides the company name and words like Home or Welcome.

Google asks for descriptive titles and is explicit about it: no “Home”.

Meta description

Present and not empty.

It's one of the sources for the snippet shown under the result.

Company structured data

The home page declares schema.org Organization or LocalBusiness.

It helps Google understand who the company is and tell it apart in results.

The reason behind each row comes from the documentation of the companies that set the rules: the pages from OpenAI, Anthropic and Perplexity on their crawlers, and Google's guides on noindex, titles and organisation data. We also measured sitemaps and product structured data, reported separately because none of these sources asks a small site for them.

Why we didn't measure llms.txt

Because Google says it doesn't use it. In its guide to generative search features, updated on 10 July 2026, the entry on llms.txt files says no special files are needed to appear in search, not even in AI features, and that creating them “will neither harm nor help” visibility, because Google Search ignores them.

“You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them.” (Google Search Central, guide to generative search features)

itria.io has an llms.txt file too, meant for other services that do read it. We mention it because the study's rule applies to us as well: we measure what search engines and assistants say they use. The same guide adds that structured data isn't a requirement for generative search, which is why we present it for what Google says it does: tell it who the company is.

The method: the ICE directory, the served HTML, the product page in a browser

The frame is ICE's public directory “Find your Italian partner”, designed for foreign companies and agents looking for Italian products or partners. The listings were collected between 19 and 31 August 2026 through searches for Puglia-related terms, such as Puglia, Apulia, Bari and Salento, and for products, such as preserves and cheese.

The directory doesn't record the region, and the search matches the text of the listings, so the companies found come from all over Italy. The listings gave 249 different websites, and we read them all, with no sampling. The sample represents the companies presenting themselves to foreign buyers in that directory; Italian food exports as a whole are beyond its reach.

For each site we read robots.txt, the home page, and the contact and product pages if linked from the home page, on 26 September 2026 between 13:34 and 14:39. We read the HTML the server sends, which is what an AI assistant that doesn't run JavaScript sees, using a tool that declares its identity, respects robots.txt and waits at least 2.5 seconds between requests.

Each measure has three possible outcomes: yes, no, and not measurable. A site that refuses the request or returns an error is counted separately and left out of the percentages. Pages, headers and robots.txt files were saved, so every measure can be rechecked against the code as it was that day.

For spec sheets we went one click further. On the evening of 26 September we opened in a browser, one site at a time, the first product page of every food website, plus up to two trade or download pages linked from the home page or the product page. Every candidate document was opened and read: general catalogues, brochures and price lists don't count.

The directory: 26 of 249 addresses lead nowhere useful

10.4% of the addresses the ICE directory shows to foreign buyers lead to something other than a company website. 14 domains no longer exist, 2 home pages return an error, and 10 pages are something else entirely: a default server page, two sites under maintenance, a WordPress install never set up, a parked domain, and a domain now taken over by an online casino.

Outcome for the 249 addressesAddressesShare
Read

195: 121 food websites across 122 addresses, 71 from other sectors, 1 unclassifiable, 1 shop under maintenance

78.3%

Home page almost empty in the HTML

14: 5 built in JavaScript, 7 that aren't really websites, 2 minimal but working

5.6%

Domain no longer exists

14

5.6%

Home page returns an error

2

0.8%

robots.txt unreachable or returning an error

14, counted as not measurable

5.6%

Access blocked or refused

6

2.4%

Other, not measurable

4: two sites that answer a browser but not the tool, a domain back online, an address mistyped in the listing

1.6%

The 10 pages that aren't websites, one by one: a server's default page, a server placeholder, a site under maintenance, an online shop under maintenance behind a password page, one being updated, a WordPress install showing the sample page with noindex switched on, a parked domain, a redirect to a domain that no longer exists, a domain now hosting an online casino, and a home page made up entirely of PHP error messages. All the addresses without a website were reopened in a browser on 26 September.

Anyone curating a directory of companies for foreign buyers could run this check in an hour: open the addresses. For the company, the website listed in a directory is a promise made to a buyer it hasn't met yet.

The table: 121 food producer websites

Of the 195 sites read, 121 belong to food and drink producers or distributors. We classified every one of them individually, reading the title, description and home page text: the directory also returns companies from other sectors, from cosmetics to machinery, because the search is text-based. Percentages are calculated on the 121 websites: two addresses in the directory lead to the same winery.

What the buyer findsWebsites with a yesOut of 121
Contact channel in the code

120

99.2%

English version

93

76.9%

Email readable by an AI assistant

107 (7 obfuscated, 7 missing from the code)

88.4%

Downloadable spec sheet, between the home page and a product page

19

15.7%

Spec sheet on the page, with the fields

10 (7 of them without a downloadable sheet)

8.3%

Terms for foreign or wholesale buyers

1 (2 counting a borderline case)

0.8%

A foreign buyer can almost always find a way to get in touch. They can download a spec sheet on one site in six, and find the terms for ordering on one in 121. Every missing piece of information turns into an email asking for it.

And response time matters: in the 2007 InsideSales.com and MIT study, companies that called back within five minutes were up to a hundred times more likely to reach the lead, as we explain in the five minutes that decide an enquiry.

The cost of the questions that arrive by email adds up to hours nobody bills for: the bill for replying by hand runs to six lines. Two documents are all it takes to join the producers who offer them: the food product spec sheet template, with its field-by-field English version, and the export price list with EXW and DDP.

Where the spec sheets are: one in two wineries, one in sixteen other producers

Downloadable spec sheets are concentrated in wine: 13 of the 25 wineries offer one (52.0%), against 6 of the other 96 food producers (6.3%). The sheet almost always sits on the individual product page, one click from the home page: 17 cases out of 19. In the other 2, it's on the page listing the wines.

Spec sheetWineries, 25Other food producers, 96
Downloadable

13 (52.0%)

6 (6.3%)

On the page, with the fields

5 (20.0%)

5 (5.2%)

At least one of the two

16 (64.0%)

10 (10.4%)

We split wineries from other food producers after the read, once we'd seen where the sheets were, using a rule written down before counting: a winery is a site whose main product is wine, and companies making both oil and wine stay with the others. The downloadable file includes the column, so the split can be rechecked.

Outside wine, the 6 downloadable sheets belong to three olive oil producers, a truffle producer, a confectioner and a dairy. Of the 19 documents, almost all are PDFs: in 3 cases the sheet is an image or a scan, and in 1 it's a wine's electronic label, a web page linked as “Technical sheet”.

Producers who have a sheet but don't show it

On 3 websites the spec sheet exists but can't be downloaded: a page of sheets behind a password, a download area that requires a login, and a “Technical sheet” link that opens a form to have it sent by email. We didn't count them as a yes, and no form was filled in.

On 4 websites the price list, prices or product list only appear after filling in a form or registering. On 11 the only product document is a general catalogue or brochure, and on 28 there is no individual product page at all: the products are all on one page or in a gallery, or not listed. Of the 27 trade or download pages read, only one offers downloadable sheets.

Search engines and AI assistants: robots.txt is open, the page says little

All 121 food websites let in Googlebot, Bingbot and the four crawlers ChatGPT, Claude and Perplexity use to build their search answers. None blocks even the training crawlers, and none uses Cloudflare's managed robots.txt, which according to its documentation only shuts out training. The door is open: what's missing is inside the page.

What search engines and assistants findWebsitesOut of 121
robots.txt open to Google and Bing

121

100%

robots.txt open to AI search crawlers

121

100%

Indexable home page, no noindex

121

100%

Page served over HTTPS

120

99.2%

Title made up only of the company name

22

18.2%

No meta description

43

35.5%

No structured data saying who the company is

53

43.8%

Product structured data on the pages read

9

7.4%

Sitemap declared or present

113

93.4%

A title like “Home - Company name” only tells the search engine and the buyer what the company is called, and Google explicitly asks sites to avoid “Home”. A home page without Organization data leaves Google to work out who the company is and to tell it apart from others with the same name. These are an afternoon's fixes, and they help all three of the site's readers.

With assistants, the game is measured in citations, because clicks rarely follow: the Pew Research Center found that people who see an AI summary on Google click a result in 8% of visits, against 15% for people who don't see one. How to work on both fronts is covered in getting found on Google and ChatGPT.

The email AI assistants can't see: 11.6%

14 of the 121 food websites have no email address an AI assistant can read: 7 obfuscate it and 7 don't include it in the code. With Cloudflare's obfuscation, for example, the address in the HTML becomes “[email protected]” and only turns back into plain text when a script runs in the browser, as Cloudflare's documentation explains. The itria.io case, and the check to run on your own site, are in our lab note on the email address AI assistants can't see.

AI assistant crawlers read the HTML exactly as it arrives: Vercel measured this across roughly 1.3 billion requests in a month, and none of the major AI crawlers runs the page's JavaScript. Ask an assistant for that company's contact details and it finds the website, but not the address.

What we corrected before publishing

A first read, on 25 September, found 2 spec sheets and 5 sets of terms for foreign buyers out of 128 sites. Checking the positive cases one by one, we found two rules that were too broad: one counted the word “Wholesale” in a menu and a retail basket's minimum order, and the other classified a vending machine manufacturer as a food business. We reread every website with the corrected rules.

MeasureFirst read, 25 SeptemberSecond read, 26 September
Food websites

128, classified by keyword

123, classified one by one

Downloadable spec sheet

2

1: the other belonged to a vending machine manufacturer

Terms for foreign buyers

5

1: the others were the word Wholesale in a menu and in a presentation, a spam text, and a retail basket

Home page almost empty

14, all put down to JavaScript

14: 5 JavaScript, 9 not really websites

Then came a third read, on the evening of 26 September. The second had looked at three pages per site and counted 1 spec sheet out of 123. Once we also opened a product page in the browser, the downloadable sheets came to 19, and rereading the websites turned up two errors in the count of food websites: all percentages are now calculated on 121.

MeasureSecond read, 26 September afternoonThird read, 26 September evening
Food websites

123

121: two addresses lead to the same winery, and an online shop under maintenance moved to the addresses without a website

Downloadable spec sheet

1, on three pages per site

19, including a product page. There were already 3 within those three pages: the rule didn't recognise “factsheet” or a sheet published as a web page

Addresses without a working website

25 out of 249

26 out of 249

The checks, and how far to trust them

Every “yes” for spec sheets and terms for foreign buyers was read individually in the saved code, and every measured site was classified individually. This case-by-case reading was done by an Itria AI agent, on the code saved on the day of the measurement, using rules written in advance. The positive cases and the addresses without a website were then reopened in a browser: this confirmed the positive cases and removed 7 addresses from the directory count, including a page that redirects to the brand's website and two sites that do respond to a browser. The borderline case is an FAQ for importers that explains what the minimum order depends on without stating it: we counted it as a no, and counting it would take the figure from 1 to 2 out of 121.

In the third read every sheet was opened and read: PDFs with a program that extracts their text, scans and images looked at one by one, and on-page sheets read in the code of the page opened in the browser, heading and fields. On 1 site the fields of the on-page sheet weren't in the code, so we counted it as a no.

robots.txt and noindex were checked on every file, and the tool correctly detects a block on a test robots.txt. Titles: every case reread individually. Descriptions, structured data and sitemaps: 20 sites drawn at random with a fixed seed and reread in the code, with the same outcome in 20 cases out of 20. With 20 cases, the true error rate for these three measures could be as high as about 15%, by the rule of three.

One limit of scope remains. For spec sheets we opened one product per site, the first listed on the product page: a site that only puts sheets on some products may come out without one, and a sheet behind restricted access isn't counted. The other measures are based on the three pages read on the afternoon of 26 September.

Downloading the data, and how to cite it

The file has one row for each of the 249 addresses, with date, outcome, sector and every measure, including those from the third read of the product page, with no company names and in shuffled order. The data is licensed under CC BY 4.0: you're free to reuse it, crediting Itria AI (itria.io) with a link to this page.

FileContentsLink
Study dataCSV, 249 rows, CC BY 4.0

Date, outcome, sector, the measures for buyers and for search engines and AI assistants, the pages read, and for food websites the product page, the downloadable, on-page or hidden sheet, and whether it's a winery.

studio-siti-produttori-alimentari-2026.csv

Questions and answers

How many food producer websites offer a downloadable spec sheet?

In Itria AI's study of 26 September 2026, 19 websites out of 121 (15.7%) offer one between the home page and a product page, and another 7 show it on the page only. Among wineries, 13 out of 25 offer one; among other food producers, 6 out of 96. The websites are those of food producers listed in the ICE directory for foreign buyers.

Every sheet was opened and read, one by one. The count looks at the first product on each site, so a site that only puts sheets on some products may come out without one.

Do food producers block Google or AI assistants in robots.txt?

No. In Itria AI's study of 26 September 2026, all 121 food websites let in Googlebot, Bingbot, OAI-SearchBot, PerplexityBot, Claude-SearchBot and Claude-User, and none of them even blocks the training crawlers.

The gap is on the page itself: in 18.2% the title is just the company name, 35.5% have no meta description and 43.8% don't use structured data to say who the company is.

Do you need an llms.txt file to appear in Google's AI answers?

Google says not. Its guide to generative search features, updated on 10 July 2026, explains that you don't need special files such as llms.txt to appear in search, AI features included, and that creating them neither helps nor harms visibility, because Google Search ignores them.

That's why Itria AI's study measures what search engines and assistants say they actually use: robots.txt, noindex, title, meta description and structured data.

Why can't an AI assistant find a company's email address on its website?

Because the address is obfuscated or isn't in the HTML at all. Assistants that read a page on the fly don't run JavaScript and only see the HTML the server sends. In Itria AI's study, 14 of the 121 food websites (11.6%) had no readable email address: 7 obfuscated it and 7 didn't include it in the code.

How was Itria AI's study of food producer websites carried out?

We read all 249 websites listed in the ICE Find your Italian partner directory entries collected through searches for Puglia-related terms and for products. For each site we read robots.txt, the home page, and the contact and product pages in the HTML as served, on 26 September 2026, respecting robots.txt and without filling in any forms. For spec sheets, we also opened the first product page in a browser.

Every site was classified one by one, and every positive case for spec sheets and terms for foreign buyers was checked individually. The data, without names, can be downloaded under CC BY 4.0.

Notes on sources

  1. ICE, Find your Italian partner: the public directory for foreign companies and agents looking for Italian products or partners. It supplies the list the study is built on: the website data are Itria's own measurements, and the study is independent of ICE.
  2. CBI, the Dutch centre that promotes sustainable production and trade between middle- and low-income countries and Europe, part of the Netherlands Enterprise Agency and funded by the Dutch Ministry of Foreign Affairs: 7 tips for finding buyers in the European processed fruit and vegetables market and 7 tips for doing business with European buyers of processed fruit and vegetables, both updated on 19 September 2023. The English sentences are quoted word for word.
  3. OpenAI, OpenAI's crawlers: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers”. robots.txt may not apply to ChatGPT-User, which is why it isn't among the measures.
  4. Anthropic, Anthropic's crawlers, updated on 7 April 2026: disabling Claude-SearchBot or Claude-User may reduce a site's visibility in user searches; ClaudeBot relates to training.
  5. Perplexity, Perplexity's crawlers: to appear in results, it recommends allowing PerplexityBot in robots.txt.
  6. Cloudflare, managed robots.txt, updated on 3 August 2026: it blocks training crawlers such as GPTBot, ClaudeBot and Google-Extended; and email address obfuscation.
  7. Vercel, The rise of the AI crawler: the major AI crawlers don't run JavaScript.
  8. Google Search Central: noindex, title links, organisation structured data (updated on 8 September 2026), sitemaps and the guide to generative search features (updated on 10 July 2026). The quotation on llms.txt is word for word.
  9. Pew Research Center, clicks with and without an AI summary, 22 July 2025: 900 US adults, 68,879 searches in March 2025, a click on a result in 8% of visits with a summary and 15% without.
  10. The measurements are by Itria AI (itria.io), taken on 26 September 2026 between 13:34 and 14:39 with a tool that reads the served HTML, respects robots.txt and declares its identity; the product pages were read in a browser on the evening and night of the same day. All the sources above were opened and read on 26 September 2026.
  11. No company names are published, either on this page or in the file. No requests were sent to the companies measured.
·The next step

See your own website the way its three readers see it.

We can take the same measurements on your website: what a buyer finds, what Google reads, what an AI assistant understands. Drop us a line about what's slowing you down: we'll make the first move, even if we never end up working together.