Crawled Content Privacy Notice
This notice explains how ArinaBot handles personal data that may appear incidentally on public web pages we crawl.
Why this is public instead of emailed to everyone
Search crawling can encounter public pages that mention many people. Contacting each person individually would involve disproportionate effort, so we make this information publicly available and provide a practical removal and objection path.
What ArinaBot does
ArinaBot visits publicly accessible web pages to build Arina's own search index, with a focus on European institutions, education, companies, public bodies, and open data. We are starting with limited, allowlisted crawl waves and expanding gradually. Crawled pages are not automatically public search results until they pass extraction, quality, safety, and indexing checks.
What personal data may be processed
- Source: publicly accessible third-party websites.
- Categories: personal data that happens to appear in public web content, such as names, public professional contact details, authorship, quotes, or institutional role information.
- We do not seek out special-category data, do not crawl logins or paywalls, and do not build behavioural profiles of individuals.
Purpose and legal basis
Purpose: to operate an independent, privacy-respecting search engine that helps people find public information.
Lawful basis: legitimate interests, GDPR Art. 6(1)(f): our interest in operating Arina and the public interest in search plurality and access to public information. We apply safeguards including robots/noindex respect, crawl limits, no profiling, no ad targeting, and removal on valid objection.
Recipients
- The public, when a crawled page becomes a public search result.
- Processors acting for us, including our EU hosting provider and
the email provider for
privacy@vaicat.com. - Competent authorities, only where legally required.
We do not sell personal data, share it with advertisers, or use it for ad targeting.
Hosting and transfers
The crawler, raw store, and search index run on EU infrastructure we control. The
email provider for privacy@vaicat.com may process correspondence outside
the EU under appropriate safeguards such as Standard Contractual Clauses or an adequacy
mechanism.
Retention
- Search index and extracted text: kept while a page remains live, relevant, and not excluded. We remove a page after two consecutive 404/410 failures.
- Raw fetched content: kept only to operate and rebuild the index, and purged on valid erasure or delisting requests.
- Operational logs: kept short-term for security and abuse handling.
- Suppression list: kept permanently so removed URLs/domains stay removed.
Your rights and removals
You have rights to access, rectification, erasure, restriction, and objection. Because we rely on legitimate interests, you may object at any time to processing concerning you. Data portability does not apply to this legitimate-interest processing. Email privacy@vaicat.com with the URL or domain concerned. On a valid request, we remove the relevant content from the public index and underlying raw storage or caches, and add the URL or domain to a permanent do-not-crawl list. We respond without undue delay and within one month.
Blocking ArinaBot
To block ArinaBot entirely, add this to your robots.txt:
User-agent: ArinaBot Disallow: /
We also honor page-level noindex and X-Robots-Tag signals.
Automated decision-making
We do not carry out automated decision-making that produces legal or similarly significant effects on individuals. Search ranking only orders results.
Changes
We will update this notice as Arina evolves and post the revised version here.