Technical focus

Disciplines we work in.

A map of the techniques, models and technologies we have a real command of. Some live inside our products. Others we research and apply in custom services.

Why we build in-house

A data analytics company depends on its ability to turn information into an asset. If that transformation comes from an external provider, you are not a data analytics company — you are a reseller with a brand.

Building in-house gives us three things no provider can give: control over the roadmap, speed to adapt the technology to real client cases, and coherence across our different lines of activity.

01
I · The raw material

Industrial-scale processing under a TDM framework

Accessing public sources at industrial scale is not trivial: it takes in-house infrastructure for authentication, lawful anti-bot handling and automated legal compliance. Everything downstream depends on getting this right.

Distributed analysis clusters

Distributed architecture across the major social networks (Facebook, X/Twitter, Instagram, TikTok, Telegram, Reddit) with load balancing, per-country user rotation and rate-limit resilience.

Native API integration

Direct connection to official APIs (Meta Graph, X/Twitter API, Reddit API) and specialised providers (RapidAPI, Apify) when it is the right tool for the case.

Lawful anti-bot for paywalled media

Integration with ScrapFly + Playwright and FlareSolverr to handle Cloudflare, DataDome, PerimeterX and Akamai over paywalled outlets — staying within the TDM legal framework and honouring opt-outs.

Centralised cookies and sessions (PupCookies)

Proprietary system for automatic cookie refresh with Puppeteer. Centralises authentication for 20+ paywalled outlets — multi-step login flows, session validation and real-time database updates.

Automated IP-law compliance

Automatic opt-out checker over robots.txt, HTTP headers and meta tags. Systematic analysis of the full media base (~130K outlets) with reservation scoring and auditability.

Lawful geolocation handling

Routing by country of origin and residential proxying where needed to reach regional sources — always inside the TDM exception and the public nature of access.

02
II · Semantic understanding

From raw information to structured intelligence

Once we have the raw material, we turn it into structured data that can be queried, filtered, aggregated and compared. This is what separates a lake of text from a business asset.

Discipline

Natural Language Processing (NLP)

Cleaning, normalisation and semantic understanding of multilingual text at industrial scale.

How we apply it

  • Language detection, tokenisation, lemmatisation and semantic deduplication of content.
  • Multilingual Named Entity Recognition (NER): people, organisations, brands, products and places.
  • Sector-aware classification with proprietary models: up to 20+ thematic categories tunable per client.
  • Canonical normalised sentiment (−100 / +100 scale), comparable across sources, languages and products.
  • Tone analysis independent from sentiment (formal, informal, aggressive, conciliatory, ironic).
  • Brand prominence (0–100): distinguishes between casual mention and deep coverage within the same text.
  • Social-network user geolocation: corrects AI errors with our own system covering 10+ countries.
Discipline

Large Language Models (LLMs)

Productive use of proprietary and open-source LLMs for large-scale summarisation, classification and rephrasing.

How we apply it

  • Selection of the right foundation LLM per use case, cost and client sensitivity. Decision made by the GeriAI architect, not tied to a single provider.
  • Prompts designed jointly with the client to guarantee sector criteria and reproducibility.
  • Per-record cost control: tracking of input/output tokens and financial reconciliation.
  • Structured outputs (JSON), parseable and persistent in database for downstream querying.
Discipline

Embeddings and semantic search

Vector representation of text and meaning-based search, not exact-match retrieval.

How we apply it

  • Scientific evaluation of embedding models (MiniLM, E5-small, MPNet, BGE, Jina) on real tasks.
  • Production model: multilingual-e5-small, chosen for measured quality/throughput balance (top1-sim 0.89).
  • Rank fusion (Reciprocal Rank Fusion) combining fuzzy and semantic retrieval.
  • Incremental indexing over corporate knowledge bases and per-client knowledge stores.
03
III · Actionable decision

Turning analysis into action

Intelligence is worthless if it does not reach the right user at the right moment. This layer turns the semantic catalogue into alerts, reports and automated actions.

Discipline

Predictive and early-alert models

Trend and event detection before things go viral or critical.

How we apply it

  • Automatic detection of emerging topics on configurable time windows.
  • Probability and impact scoring on detected events.
  • Brand-prominence and voice-prominence models over the analysed content.
  • Real-time alerts delivered through Telegram, email and other channels.
GeriAI · Mochis

Autonomous agents

Agents that reason over context, decide what matters and write the message themselves. Not rules: reasoning.

See GeriAI (cognitive layer) →

How we apply it

  • Automatic identification of the most relevant topics in each period over the semantic catalogue.
  • Generation of editorial drafts and narrative reports from the underlying analysis.
  • Domain-specialised agents: surveys, elections, PR, sector observatories.
  • Conversational agents and knowledge bots over proprietary or client knowledge bases.
  • Natural-language orchestrator: automatic provisioning of services from free-text briefings — hours to minutes.
04
IV · Data infrastructure

The layer that holds everything up

A hybrid operational/analytical architecture running 24/7, sized for the real volumes of the public universe of the Internet.

Discipline

Big Data pipelines and Elasticsearch

24/7 batch infrastructure underpinning everything else. Operational MySQL + analytical BigQuery + Elasticsearch for real-time retrieval.

How we apply it

  • Processing, analysis and enrichment pipelines across hundreds of sources and integrated APIs.
  • Four Elasticsearch clusters split by thematic domain (Press, DG, PRO, Social) with monthly time-based indices.
  • Production volumes: in the order of 3–7M analysed digital-press publications and 300K–4M social-network signals processed per month.
  • Cross-schema synchronisation between operational MySQL and analytical BigQuery for history and modelling.
  • Industrial resilience: retries, throttling and regular delivery with no losses within the usual window.
05
V · Value metrics

From data to economic impact

What our infrastructure measures does not stop at technical numbers: it translates into advertising value, media tiering and comparable impact metrics.

Discipline

Automated media enrichment (MediaAudit)

Proprietary ETL pipeline keeping a cross-source catalogue of global digital media continuously up to date.

How we apply it

  • More than 10,000 outlets catalogued with audience, ranking and geographic distribution metadata.
  • Cross-ranking with SEMrush, SimilarWeb and Moz for a comparable multi-source view.
  • Ad-revenue estimation and automatic tier classification by relevance.
  • Automatic synchronisation with internal products and with clients who need to segment their media universe.
Discipline

Automatic advertising value (AdValue)

Calculation of the economic equivalent of every mention, per channel, with differentiated formulas and market data.

How we apply it

  • Analogue media (print, radio, TV): official rate cards by surface or screen-time, VAT included.
  • Social networks: (impressions × CPM) / 1000, with platform-specific CPMs (TikTok, YouTube, Instagram, X, Facebook).
  • Digital media: estimated daily ad revenue of the outlet × view share of the piece.
  • Output in comparable USD across channels — a direct basis for ROI and communications-spend justification.
Operations

Responsible analysis practices

Rules of engagement we apply to every analysis, whatever the sector and language.

How we apply it

  • Analysis of lawful sources from the public universe of the Internet.
  • Respect for machine-readable rights reservations (robots.txt, headers, standardised metadata).
  • Automated opt-out checker running continuously over our full media base.
  • Transformed nature of the delivered output: ~80–90% processing relative to the original source.
Research

Where we are looking next

Beyond what already runs in production, we keep open research lines on the directions that will define the next generation of data-based products.

FUTURO / RESEARCH

Model Context Protocol (MCP)

Exposing our semantic catalogue as an MCP server (Anthropic standard), so that external LLMs can query our infrastructure directly with no bespoke integration.

See GeriAI →
FUTURO / RESEARCH

New categories of data-based product

Internal research lines on new products that leverage the public universe of the Internet in ways the current market does not cover.

Propose a case →
Let's talk

Need something custom?

If your case does not fit a standard product, we deliver custom development services on data and AI: data lakes, integrations, sector-specific models, dedicated dashboards. Let's talk.

Talk to the team

Response in under 24h · SAT + Sales