Practical Machine Learning · Open Science · Modern Engineering

We build language models and AI systems that actually work.

We bridge the gap between NLP research and solid software engineering. We train models for Polish, build open-source tools, and openly share what we learn along the way.

About us

From research and experiments to production

We solve real-world engineering challenges with machine learning. We cover the entire journey: from curating and cleaning datasets, through model architecture and training, to reliable deployments in production environments. When off-the-shelf tools fall short, we design and build what's needed from the ground up.

A central focus of our work is the Polish language. With its rich inflection, complex grammar, and limited public datasets, Polish demands much more than blindly querying third-party APIs — it requires a deep understanding of methods and what actually happens under the hood.

We emphasize data sovereignty and on-premise deployments, including air-gapped environments once models and dependencies are provisioned locally. Local processing and data masking support GDPR / RODO and EU AI Act requirements without requiring content to be sent to external APIs. Compliance also depends on deployment configuration, use case, and organizational procedures — the tools alone do not guarantee it.

Our solutions

Tools and systems ready to use

LLM Router

A high-performance open-source AI gateway uniting local engines with cloud APIs. Delivers query routing, data protection, and load balancing.

Details & architecture

Radar Informacji

Real-time topic detection and trend monitoring across incoming text streams, powered by semantic similarity and source clustering.

Details & algorithm

PII Masker

Fast and reliable anonymization of personal data and confidential information before sending text to external language models.

Details & privacy engine

RDL Playground AI

An AI testing ground with a news summary stream, articles generated from questions, and daily information overviews. Runs locally, with no account required.

Details & modules

Tech stack

The technology foundation behind our systems

We combine language models, semantic search, and data protection tools. This stack supports our own solutions — from local news analysis to managing traffic between models.

PyTorch model training & fine-tuning
Transformers language models & classifiers
Sentence-Transformers text embeddings
vLLM language model inference
Ollama local model serving
Milvus vector database & semantic search
Redis traffic coordination & rate limiting
ONNX NER model inference on CPU

Models & training

We develop dedicated models for Polish: generative pLLama models, semantic encoders, extractive QA, and NER classifiers. Built with PyTorch and the Hugging Face ecosystem, optimized for high throughput on local hardware (RDL Playground AI).

Inference & LLM routing

The foundation of LLM Router: connecting local inference engines (vLLM, Ollama) with cloud APIs via OpenAI/Anthropic endpoints. Features configurable plugin pipelines, advanced load-balancing strategies, and Redis coordination for self-hosted sovereignty.

Search & information analysis

The analytical core of Radar Informacji: semantic search via Milvus, dense text embeddings, non-linear dimensionality reduction (t-SNE/UMAP), and dynamic HDBSCAN clustering for uncategorized trend discovery in news streams.

Anonymisation & data protection

The engine behind PII Masker: dual-layer privacy protection (deterministic FastMasker with checksum validation + RoBERTa NER). ONNX INT8 quantization enables masking on CPU, supporting data minimization in projects subject to GDPR and the EU AI Act.

Open source & resources

Code, models, and research we share

We believe open science accelerates innovation. We publish our core libraries, model weights, curated datasets, and engineering findings for the global AI community.

GitHub

Open-source repositories for our AI tools, routers, and utility libraries, ready for exploration and self-hosting.

radlab-dev-group

Language models

Polish-tailored models from the pLLama family (1B to 70B), semantic encoders, and extractive QA checkpoints.

huggingface.co/radlab

Datasets

Curated and cleaned Polish text datasets, including polish-sts, legal-mc4-pl, and specialized domain corpora.

browse datasets

Engineering blog

In-depth insights into training dynamics, architecture trade-offs, and practical lessons from our ML experiments.

read the blog

Contact & Collaboration

Have questions or want to collaborate?

Reach out by email. We are always glad to consult on system architecture, production deployments, or the specifics of natural language processing for Polish.

  • AI Gateway Architecture & Cost Optimization — LLM Router deployments, vLLM / Ollama engine pooling, advanced load-balancing strategies, and API cost reduction.
  • Fine-tuning & Evaluation for Polish — Domain-specific model fine-tuning, benchmark evaluation, and quantization for self-hosted infrastructure.
  • On-Premise Knowledge Bases, RAG & Privacy — Secure enterprise RAG systems, vector search integration with Milvus, and PII masking for GDPR compliance.

Blog

Latest from the blog

All posts