The pilot never reaches production
The demo works, but quality, permissions, cost or operations are still unresolved.
01 / AI DEVELOPER & LLM FREELANCER BERLIN
I build AI functions for real specialist processes, from data and evaluation pipelines through RAG to controlled agents with tool access. Quality, permissions, cost and operations become measurable before production use.
THE PROJECT QUESTION
Start with a tightly scoped use case, representative test data and an architecture that controls the model, data access and tool permissions separately. Measured results determine whether the system should be expanded.
WHEN I SUPPORT YOUR TEAM
The first decision is not which model to buy. We clarify the task, data, acceptable failure modes and success criteria. That makes it clear whether RAG, an agent, conventional automation or a combination is commercially useful.
The demo works, but quality, permissions, cost or operations are still unresolved.
Documents and specialist data need to become reliably accessible through RAG, with sources and clear access rules.
Legal, technical or internal data cannot be sent to external model providers without control.
Your developers know the product; the missing experience is retrieval, agents, evals and local operation.
AI ENGINEERING / DECISION SYSTEM
A dependable AI system does not emerge from selecting a famous model. Data quality, retrieval, model behaviour, quantisation and operations are designed separately and compared on representative tasks. Investment goes into the component that demonstrably needs improvement.
Professional success criteria, difficult cases and acceptable failure modes come before architecture. A simple baseline shows whether RAG, an agent, fine-tuning or conventional software creates measurable value.
I build pipelines for collection, cleaning, deduplication, normalisation, annotation, versioning and quality assurance. Training, validation and test data remain separate so overfitting cannot masquerade as progress.
Document processing, chunking, metadata, permission filters, embedding model, vector search and reranking are evaluated separately. This reveals whether retrieval or answer generation is the actual source of error.
Frontier, local and specialised smaller models are compared on the same tasks. Qwen3.6-27B and Gemma 4 31B are possible local candidates, but latency, quality, privacy and operations decide, not a public benchmark.
FP16, BF16 and quantised variants can change memory use, context, concurrency and answer quality. I measure those effects on the available hardware instead of treating quantisation as a free optimisation.
Fine-tuning does not repair poor or contradictory training data. Only after prompting, RAG and tool logic have been evaluated and a stable behavioural gap remains do I test SFT or adapter tuning against an independent regression suite.
Quality, privacy, latency, throughput and operating cost are measured together. The architecture remains interchangeable when better models become available.
SELECTED PROJECTS / 2
Real work, explained clearly. Confidential interfaces and data remain consistently anonymised.
SCREEN / 01Anonymised UI reconstruction · no client or case data
A platform for a legal-sector company makes extensive legal sources accessible, structures matters and prepares working drafts. Several lawyers supplied real test cases and evaluated professional quality.
SCREEN / 02Anonymised agent console · fully synthetic target data
An autonomous system plans tests, uses isolated security tools, evaluates evidence and documents reproducible vulnerabilities. The architecture works with different model classes.
AI-ASSISTED DEVELOPMENT
AI can accelerate analysis, implementation, tests and documentation. Model choice follows the task, data class and required level of control.
For non-sensitive data, powerful models accelerate research, implementation, tests and documentation. Results are reviewed like every other contribution.
Models can run locally or inside your controlled infrastructure for confidential code, legal data or internal documents.
CLIENT FEEDBACK / AI DEVELOPMENT
“Damian treats the client’s problem as a challenge and the result is always outstanding. He goes deeply into every problem until he finds a solution. Giving up is not an option for Damian.
“Damian is a real stroke of luck. Clear, open communication and completely transparent work make the collaboration a success. Our requirements were implemented competently and flexibly.
FREQUENTLY ASKED QUESTIONS / AI DEVELOPER & LLM FREELANCER BERLIN
Clear answers about the start, collaboration, data protection and budget before you invest time in a long sales process.
Ask your question directly ↗Yes. I can own an AI component, lead architecture and evaluation, or support your team during implementation. Decisions and interfaces are documented so knowledge remains inside the company.
No. Depending on the required level of protection, we can use local models, self-hosted inference or clearly separated hybrid workflows. External models are used only where the data class and agreed rules allow it.
Representative tasks and success criteria are defined before implementation. We then measure source grounding, task success, errors, cost and policy violations with repeatable tests.
Yes, if the first use case is narrow. A focused pilot answers the riskiest questions first and avoids spending money on a technically impressive but professionally useless demo.
Frontier models accelerate analysis, implementation, tests and documentation. That speed is combined with manual reviews and automated tests. Local models are available for sensitive work.
Data quality, integrations and quality requirements matter more than the model name. After a short assessment, you receive a sensible first scope with assumptions and a budget range. A focused feasibility review or pilot limits risk before larger delivery.
A workstation with 256 GB RAM and two RTX 4090 GPUs is available for compact and medium-sized models. Two DGX Spark systems extend local memory for larger models and contexts. Candidates such as Qwen3.6-27B and Gemma 4 31B are compared at suitable quantisation levels using your tasks. Quality, latency, throughput and privacy decide, not a public leaderboard.
When a narrow, recurring behavioural problem cannot be solved reliably through prompting, RAG or tool logic and enough high-quality training data exists. Without clean data, a separate test set and a baseline, fine-tuning can preserve errors instead of improving quality.
A tightly scoped pilot can often be evaluated within a few weeks. The decisive factors are test data, integrations and the quality threshold. The expected result, stop criteria and next decision point are agreed before work starts.
THREE DISCIPLINES, ONE PARTNER
Ambitious projects benefit when modern AI, scalable product development and offensive security are considered together.