View all jobs

Data Scientist (pharmaceutical launch planning) - September 2026

  • remote, remote

For one of our clients in the pharma industry we are looking for a Data Scientist (pharmaceutical launch planning)

Project name: Generate Insights from Hidden Knowledge (GIHK)

Project description:
The services are requested as part of the GIHK project; the Project has the goal of transforming knowledge graphs, RAG pipelines, and enterprise data into AI-driven personas that support pharmaceutical launch planning, customer experience, and evidence-based decision-making.

Background to the assignment:
The project requires a combination of strong software engineering skills with hands-on experience in LLM applications, agentic AI workflows, RAG, knowledge graphs, and scalable data platforms, because this expertise is not available internally the external contractor has a unique position compared to the client's internal project staff and provides significantly different services than the internal staff.

Tasks:
  • Technical Development of Synthetic AI persona using knowledge graphs, RAG, GraphRAG, chunking strategies, and enterprise data sources.
  • Technical Design and implementation LLM-based features, agentic workflows, and AI-driven insight generation capabilities.
  • Technical development and build of a scalable backend services, data pipelines, and integrations for AI/ML and persona-related use cases.
  • Technical translation of business and product requirements into technical concepts, user stories, and working software.
  • Test, refine, and optimize AI components for quality, performance, scalability, and reliability.
Deliverables: Creation of comprehensive documentation with all results regarding the above mentioned tasks with subsequent handover to client for review and approval for further usage.

Key Technologies:
  • Azure AI components, Azure OpenAI, Amazon Bedrock, OpenAI, Gemini, and Anthropic
  • Agentic workflows, LangChain, LangFuse, Haystack, prompt orchestration, tracing, and evaluation
  • RAG, GraphRAG, embeddings, vector search, semantic chunking, knowledge graphs
  • Snowflake and enterprise data platforms for structured and semi-structured data
  • Python, Rust, TypeScript, Node.js, FastAPI, Java, PostgreSQL
  • Docker, Openshift, Kubernetes and CI/CD

Required Qualifications:
  • Strong professional experience in software engineering, preferably with production-grade AI/ML or data-driven platforms.
  • Hands-on experience with LLM applications, RAG pipelines, AI orchestration, and modern backend development.
  • Good understanding of knowledge graphs, vector search, embeddings, chunking strategies, and unstructured data processing.
  • Experience with cloud-native development, APIs, microservices, testing, CI/CD, and scalable system design.
  • Ability to communicate complex technical topics clearly to both technical and non-technical stakeholders.

Preferred Skills:
  • Experience with LangChain, LangFuse, agentic AI tooling, prompt tracing, evaluation, and observability.
  • Experience with Azure AI components, Amazon Bedrock, Snowflake, ArangoDB, and multi-provider LLM orchestration.
  • Experience with OCR, document intelligence, data extraction, or transformation of unstructured content into structured knowledge assets.
  • Experience with document chunking strategies, semantic chunking, metadata enrichment, retrieval optimization, and improving context quality for RAG-based systems.

Nice to have:
  • Experience in pharma, healthcare, biomedical data, launch planning, or customer experience use cases.

Start: ASAP
Capacity: full-time, 40h/week
Duration: till end of February 2027
Location: remote
Hourly rate: 70-75 Euro/h