Istella’s data assets form the informational foundation of the AI platform: a cognitive infrastructure that integrates internal and external data to enrich organizational knowledge, while respecting ownership, control, and protection of corporate information.

Istella collects, organizes, and leverages heterogeneous information assets, such as web data, multimedia content, social information, and specialized datasets, transforming them into a structured foundation for search, analysis, and artificial intelligence applications. These assets can be integrated with an organization’s proprietary data, documents, and sources, keeping informational domains separate and preserving the customer’s control over their own knowledge assets. The result is a richer, more articulated context in which the platform can identify connections, retrieve relevant information, and support retrieval, reasoning, and content generation activities in enterprise scenarios.

Qualified, selected, and defined data becomes cognitive infrastructure: the starting point for fueling models, enriching knowledge, and making the AI ​​platform more precise, contextual, and useful.

Web data

Istella gathers online information and transforms it into knowledge that enriches the platform’s data. Sources are processed continuously, ensuring an always-up-to-date foundation ready to support search, analysis, and AI applications.

Istella collects, analyzes, enriches, and indexes 8 billion URLs, with over 15 billion URLs discovered and 150 million URLs updated daily. Approximately 400 signals are extracted from each page, while 2,000 RSS feeds and 1,500 news sites are analyzed on an ongoing basis, every 10 minutes.

The value of this asset lies not only in scale, but in its ability to provide contextual understanding that goes beyond the customer’s proprietary data. The web thus becomes a living source of up-to-date information, useful for connecting diverse sources, retrieving relevant elements, and delivering more precise and contextualized answers.

Multimedia

Istella’s multimedia heritage expands the boundaries of knowledge beyond text, integrating visual and audiovisual content into the platform that becomes part of its informational intelligence. The current base includes 16 terabytes of indexed images and over 140 million indexed videos, with a daily inflow of approximately 10,000 new images and 170,000 new videos.

The distinctive value of this asset lies in its ability to link multimedia content to documents, web signals, and the knowledge graph, deepening analytical capabilities and improving the quality of responses produced by the platform. In this way, multimedia becomes a structural component of the data heritage, designed to support advanced search and multimodal AI applications.

Social

Social data also enrich Istella’s information heritage with a real-time view of digital conversations. The system indexes over 2.3 billion tweets and collects approximately 5 million tweets daily, with access also to Facebook and Instagram firehoses, integrating signals that reflect events, opinions, and emerging dynamics.

This base is not only for observing the present, but also for identifying emerging trends, recurring themes, and relationships among people, brands, places, and events. In this way, social data fuels semantic enrichment, strengthens the knowledge graph, and supports AI applications that must interpret live, fast-moving, and constantly evolving signals.

Alongside this dynamic information heritage are public datasets that testify to the scientific and technological expertise Istella has developed in the fields of research, ranking, and AI model evaluation.

Public datasets and expertise

Istella22 and LETOR are public benchmarks developed for research and experimentation. Their value for Istella lies in their ability to document proprietary expertise in building ranking systems, representing query-document relationships, and evaluating result relevance. The datasets are subject to specific licensing conditions and intended for non-commercial purposes. For information on access, usage, and citation requirements, please contact Istella.

Istella22

Istella22 is the dataset through which Istella demonstrates its expertise in ranking and large-scale search model evaluation. The dataset was designed to enable rigorous comparison between traditional learning-to-rank approaches and neural models, providing both query and document text and industrial-grade structured ranking features in a single environment.

The corpus comprises 8.4 million web documents, 220 query-document features, 2,198 test queries, and 10,693 relevance judgments distributed across a 5-level scale. This combination makes Istella22 particularly suitable for comparative evaluation of ranking, re-ranking, semantic retrieval, and AI models applied to search.

The dataset’s strength lies in its ability to combine, within the same benchmark, textual signals and structured signals, allowing consistent measurement of the effectiveness of both classical and neural techniques. For this reason, Istella22 represents an important reference and a concrete demonstration of the maturity achieved in building infrastructure for ranking and content understanding.

LETOR Dataset

The LETOR dataset is one of the main public benchmarks developed by Istella for Learning to Rank: a dataset designed to train and evaluate models that order search results by relevance, testing precision, efficiency, and ability to scale across large volumes of queries and documents.

The full version includes 33,018 queries, 220 features per query-document pair, and 10,454,629 examples annotated with relevance judgments from 0 to 4. The dataset is split into training and test sets according to an 80%-20% scheme, with an average of 316 examples per query, making it particularly suitable for evaluating feature-based ranking algorithms. In addition to the full version, Istella also offers Istella-S LETOR, a more compact version with 3,408,630 pairs and approximately 103 examples per query, and Istella-X LETOR, an extended version with 26,791,447 pairs obtained by retrieving up to 5,000 documents per query according to BM25F ranking.

The LETOR family documents Istella’s ability to work with large volumes of queries and documents, design ranking signals, and evaluate result quality in high-information-intensity environments.