Alex Soto & Markus Eisele
RAG like a hero with Docling
#1about 3 minutes
Using RAG to enrich LLMs with proprietary data
Retrieval-augmented generation (RAG) is the key to making large language models useful for enterprises by providing them with up-to-date, proprietary information.
#2about 4 minutes
The challenge of parsing complex document structures
Simple document parsers can misinterpret layouts like multi-column text, leading to corrupted data and incorrect outputs from the language model.
#3about 3 minutes
Using Docling to convert documents into structured formats
Docling is an open-source tool that acts like an advanced OCR service, converting various binary document formats into a structured, parsable tree.
#4about 7 minutes
Demo of a basic RAG ingestion pipeline
A live demonstration shows how a Quarkus application uses Docling to ingest a PDF, generate embeddings, and store the resulting chunks and vectors in Redis.
#5about 3 minutes
Securing RAG against data poisoning and leaks
To prevent data poisoning and sensitive data leaks, it is crucial to sanitize documents, verify their signatures, and use tools for PII masking.
#6about 4 minutes
Mitigating vector store attacks and encryption challenges
Vector stores are vulnerable to attacks like close vector modification and reversal, and standard encryption breaks vector distance, requiring specialized solutions.
#7about 5 minutes
Demo of a secure ingestion pipeline in action
A final demonstration showcases a secure pipeline that verifies document signatures, anonymizes sensitive data, and encrypts vectors before storing them.
Related jobs
Jobs that call for the skills explored in this talk.
ROSEN Technology and Research Center GmbH
Osnabrück, Germany
Senior
TypeScript
React
+3
Wilken GmbH
Ulm, Germany
Senior
Kubernetes
AI Frameworks
+3
VECTOR Informatik
Stuttgart, Germany
Senior
Java
IT Security
Matching moments
07:39 MIN
Prompt injection as an unsolved AI security problem
AI in the Open and in Browsers - Tarek Ziadé
01:15 MIN
Crypto crime, EU regulation, and working while you sleep
Fake or News: Self-Driving Cars on Subscription, Crypto Attacks Rising and Working While You Sleep - Théodore Lefèvre
04:57 MIN
Increasing the value of talk recordings post-event
Cat Herding with Lions and Tigers - Christian Heilmann
02:49 MIN
Using AI to overcome challenges in systems programming
AI in the Open and in Browsers - Tarek Ziadé
01:06 MIN
Malware campaigns, cloud latency, and government IT theft
Fake or News: Self-Driving Cars on Subscription, Crypto Attacks Rising and Working While You Sleep - Théodore Lefèvre
08:29 MIN
How AI threatens the open source documentation business model
WeAreDevelopers LIVE – AI, Freelancing, Keeping Up with Tech and More
05:55 MIN
The security risks of AI-generated code and slopsquatting
Slopquatting, API Keys, Fun with Fonts, Recruiters vs AI and more - The Best of LIVE 2025 - Part 2
06:28 MIN
Using AI agents to modernize legacy COBOL systems
Devs vs. Marketers, COBOL and Copilot, Make Live Coding Easy and more - The Best of LIVE 2025 - Part 3
Featured Partners
Related Videos
Carl Lapierre - Exploring Advanced Patterns in Retrieval-Augmented Generation
Carl Lapierre
Building Blocks of RAG: From Understanding to Implementation
Ashish Sharma
Accelerating GenAI Development: Harnessing Astra DB Vector Store and Langflow for LLM-Powered Apps
Dieter Flick & Michel de Ru
Build RAG from Scratch
Phil Nash
Large Language Models ❤️ Knowledge Graphs
Michael Hunger
Beyond the Hype: Building Trustworthy and Reliable LLM Applications with Guardrails
Alex Soto
Building AI Applications with LangChain and Node.js
Julián Duque
Langchain4J - An Introduction for Impatient Developers
Juarez Junior
Related Articles
View all articles



From learning to earning
Jobs that call for the skills explored in this talk.

Forschungszentrum Jülich GmbH
Jülich, Germany
Intermediate
Senior
Linux
Docker
AI Frameworks
Machine Learning

Riverty GmbH
Verl, Germany
Remote
Java
Python
TypeScript

Riverty GmbH
Berlin, Germany
Remote
Java
Python
TypeScript

The Rolewe
Charing Cross, United Kingdom
API
Python
Machine Learning



Robert Ragge GmbH
Senior
API
Python
Terraform
Kubernetes
A/B testing
+3

