Data scientist – Synthetic Data Generation

  • Clever Access/ InovIntell -
  • Tunis, Tunisie
  • Il'y a 1 mois
Postes vacants:
1 poste ouvert
Type d'emploi désiré :
CDI, Temps plein
Experience :
3 à 5 ans
Niveau d'étude :
Ingénieur
Langue :
Anglais
Genre :
Indifférent

Description de l'emploi

About InovIntell

InovIntell is a startup developing AI-empowered solutions for life-sciences companies. Our solutions are used at all stages of the development and evaluation process of medical products. The company relies on an international team of experts in artificial intelligence, data science, software engineering and health economics.

For more information, please visit www.inovintell.com.

Position Summary

One of our core workstreams is the development and application of generative models that produce synthetic patient data. Synthetic data reproduce the statistical properties of and relationships in real patient-level data, and are used for anonymisation, missing-data imputation, extrapolation, counterfactual simulation and indirect treatment comparisons across populations that are otherwise inaccessible with standard methods.

We are seeking a talented and motivated Data Scientist to strengthen this team. The successful candidate will contribute to the development, validation and application of generative models for tabular and longitudinal patient data, in collaboration with other data scientists, machine-learning engineers, statisticians and health-economics experts.

This research-oriented role includes a strong methodological component and may offer the opportunity to join a PhD programme in parallel. 

Key Responsibilities

  • Write clean, well-tested Python code for the training, validation and application of generative machine-learning models for tabular and longitudinal patient data (e.g. VAE- and neural-ODE-based architectures).
  • Build and maintain reproducible pipelines for data preprocessing, model training, hyperparameter optimisation and evaluation.
  • Analyse and preprocess ground-truth clinical data, including handling of missing data and censoring.
  • Implement and adapt ML methods for population matching (e.g. reweighting, optimal transport).
  • Design and implement validation metrics for synthetic data and quantify uncertainty in model outputs.
  • Propose adaptations to model architectures, training algorithms and validation methods for new applications, and implement improvements to existing methods.
  • Collaborate proactively with internal and external stakeholders to translate methodological and business requirements into working code.
  • Document methods clearly for internal handoff and for inclusion in analysis plans and scientific reports.
  • Continuously evaluate state-of-the-art generative and survival-modelling methods to improve our offering. 

    Why InovIntell?

    • Be part of pioneering projects that drive methodological advances in the life-sciences sector.
    • Join a human-sized, international, multidisciplinary team.
    • A dynamic environment with room for fast professional growth.
    • Flexible working conditions that respect work-life balance.

    Application Process

    Interested candidates are invited to submit their CV and a cover letter in English to [email protected], with the subject line “Data scientist / SDG”. In your cover letter, please describe an ML project that you have led in a professional setting and that you are proud of.

     

    Join us at InovIntell, where your expertise will help shape the future of AI in life sciences. Apply today!

Exigences de l'emploi

Requirements

Essential

  • Proven experience in machine learning and deep learning, including the development of generative models (e.g. VAEs, neural ODEs, diffusion models or GANs), ideally for tabular or longitudinal data.
  • Excellent programming skills in Python and the core ML/scientific stack (PyTorch and/or TensorFlow, NumPy, SciPy, scikit-learn).
  • Solid software-engineering practice: version control (Git), automated testing, containerisation (Docker) and reproducible workflows.
  • Experience preprocessing structured/tabular data, including handling of missing data.
  • Strong applied-statistics foundation (probability, regression) and the ability to reason about relationships between model inputs and outputs.
  • Strong analytical and problem-solving skills, and a proactive approach to learning new methods and technologies.
  • Excellent written and verbal communication in English, and the ability to work effectively in a hybrid, international team.

Desirable

  • Familiarity with optimal transport methods.
  • Familiarity with statistical analysis of healthcare data, indirect treatment comparisons and causal inference methods.
  • Experience with hyperparameter-optimisation frameworks (e.g. Optuna) and with high performance computing or distributed training.
  • Bonus points for candidates with experience in the life-sciences, health-economics or clinical-trial-data setting.

Date d'expiration

13/08/2026