What is omgili? AI crawler guide
omgili by Webz.io: Webz.io crawler for collecting web data sold through APIs and datasets. Check its reported user-agent, robots.txt behavior, source, and verification guidance.
Webz.io crawler for collecting web data sold through APIs and datasets.
What is omgili?
omgili is a web crawler operated by Webz.io that collects publicly available web data for inclusion in commercial APIs and datasets. It identifies itself with the user-agent token omgili and respects the Robots Exclusion Protocol. The crawler gathers content from across the web, which Webz.io then structures and sells to businesses, researchers, and developers for use in monitoring, analytics, and machine learning applications. Its activity is part of Webz.io's broader data-as-a-service offering, turning unstructured web information into structured feeds.
What it's for
If omgili crawls your site, your content may be extracted and incorporated into Webz.io's data products. This means your public pages could appear in downstream monitoring dashboards, research datasets, or AI training corpora sold to third parties. The crawler's visits can influence how your information is distributed and monetized beyond your own channels, potentially affecting your brand's reach and the control you have over your content's commercial use.
Allowing omgili only creates the possibility of retrieval. To find out whether your pages are selected, measure which answer engines actually cite your pages using a stable query set.
How to handle omgili
To prevent omgili from crawling your site, add a robots.txt rule that disallows the user-agent token omgili for the paths you want to protect. The crawler honors standard robots.txt directives, so blocking it will stop collection of the specified pages. If you do not block it, omgili will access and harvest publicly available content according to its crawl policies.
robots.txt rule
User-agent: omgili Disallow: /
Blocking cost
Blocking omgili may prevent your content from appearing in Webz.io's data feeds, which could reduce your visibility in AI training data, research datasets, and monitoring tools that rely on their services.
Examples
- A news website's articles are crawled by omgili and later appear in a Webz.io dataset used for media monitoring.
- A corporate blog post is collected by omgili and ends up in a machine learning training corpus sold by Webz.io.
- An e-commerce product page is harvested by omgili and included in a competitive pricing analysis feed provided by Webz.io.
Related bots
- Brightbot: Also tracked as a general crawler.
- Amazonbot: Also tracked as a general crawler.
- ImagesiftBot: Also tracked as a general crawler.
- omgilibot: Another Webz.io general crawler to compare.
- Diffbot: Also tracked as a general crawler.
- GoogleOther: Also tracked as a general crawler.
- GoogleOther-Image: Also tracked as a general crawler.
- GoogleOther-Video: Also tracked as a general crawler.
- Panscient: Also tracked as a general crawler.
- Robots.txt: Robots.txt is the control file used to allow or block omgili.
- AI Crawlers: omgili is a concrete crawler example for this concept.
Frequently Asked Questions
Who operates the omgili crawler?
The omgili crawler is operated by Webz.io, a company that provides web data as a service through APIs and datasets.
Does omgili respect robots.txt?
Yes, omgili honors the Robots Exclusion Protocol. You can control its access to your site by setting rules in your robots.txt file.
What happens to the data collected by omgili?
Data collected by omgili is processed and sold by Webz.io as part of their commercial data products, which may be used for monitoring, research, and AI training.
How can I stop omgili from crawling my site?
You can block omgili by adding a Disallow rule for the user-agent omgili in your robots.txt file. The crawler will comply with the directive.
Will blocking omgili affect my site's visibility in other services?
Blocking omgili may prevent your content from being included in Webz.io's data feeds, which could reduce your presence in downstream applications that rely on their data, such as AI models or monitoring tools.
Data & Sources
- Webz.io documentation - Primary source for omgili crawler details.