Job Description
DIGIPOL is a company that exists to modernize traditional industries and outdated workflows, driving real, sector-wide transformation. We help established businesses enter and thrive in e-commerce, while also pushing digital-native players to perform even better. Our work cuts deep: accelerating innovation, lifting customer experiences, and shoring up business fundamentals. To get there, we keep hierarchy flat, teams small, and talent dense. We move with the ambition and speed of a startup, and we treat AI as a working partner, so human insight and machine intelligence constantly sharpen each other.
The Opportunity
As an ML Scientist focused on post-training, you will improve the capabilities and behavior of small language models for the Farsi language. You will own the full experimental loop: identifying model weaknesses, developing data and training strategies, running controlled experiments, evaluating results, and turning successful ideas into high-quality checkpoints and reusable methods.
You will work closely with DIGIPOL’s research and infrastructure teams. Signals from real deployments will give you concrete, difficult problems to solve; your work will translate those signals into broader improvements in model quality, reliability, and usefulness.
This is a role for a scientist who builds. Strong ideas matter, but so do careful data work, reliable implementations, and measurable improvements in working models.
What We’re Looking For
We need someone who:
- Thinks experimentally: You form clear hypotheses, establish strong baselines, design informative ablations, and let evidence guide the next step.
- Understands models through data: You can identify gaps in training data and design curation, filtering, generation, and supervision strategies to address them.
- Owns the full post-training loop: You are comfortable moving between datasets, training, evaluation, error analysis, and implementation.
- Builds with rigor: Your experiments are reproducible, your conclusions are well supported, and successful research can be integrated into shared systems.
- Collaborates across boundaries: You communicate clearly with research, engineering, and applied teams.
The Work
- Design and execute post-training strategies for language and multimodal models, including supervised fine-tuning, preference optimization, reinforcement learning, and distillation.
- Build and curate high-quality training data using human, synthetic, and model-generated signals, with particular attention to Farsi-language and domain-specific capabilities.
- Develop evaluations that expose meaningful capability and reliability gaps across Farsi and global use cases.
- Conduct systematic error analysis and use the results to improve data mixtures, objectives, training methods, and model behavior.
- Run controlled experiments and ablations, interpret results, and communicate clear recommendations.
- Develop reliable, scalable training and evaluation pipelines in collaboration with model infrastructure teams.
- Contribute methods, tooling, datasets, and findings that accelerate post-training work.
Desired Experience
Must-have:
- Hands-on experience post-training small language models or multimodal models.
- Strong understanding of fine-tuning fundamentals and current post-training and RL methods.
- Solid engineering skills and proficiency with the open-source ML ecosystem.
- Experience building, curating, or assessing training and evaluation data.
- Ability to turn research ideas into reliable implementations and measurable model improvements.
- Proficiency in English, including the ability to collaborate on complex technical work with other teams.
- Experience leveraging agents to amplify your own work.
Nice-to-have:
- Experience with SFT, DPO, or RL methods for foundation models.
- Experience post-training multimodal models involving text, vision, or audio.
- Experience developing synthetic data pipelines, reward models, verifiers, or model-based evaluations.
What Success Looks Like
- You deliver checkpoints with clear, measurable improvements on important Farsi and domain-specific capabilities.
- You establish evaluations and error-analysis practices that reliably identify model weaknesses and guide post-training priorities.
- You develop data or training methods that become part of our repeatable post-training workflow.
- You run high-quality experiments at increasing scale and turn the results into clear technical decisions.
- Your work materially influences our model roadmap and improves systems used in real deployments.
What we offer
- Incredibly talented, entrepreneurial teams in small, result-oriented squads.
- An exceptional opportunity for growth and unlimited backing for learning.
- Flexible hours, remote working.
Commitment & contract
Permanent or fixed-term. Part-time.
The selection process
We prioritize verifiable excellence and demonstrated ability over seniority or a flawless CV. So if you like the role and believe you can grow into it, don’t self-reject. Degrees and publications aren’t required, and we encourage you to apply even if your experience doesn’t check every box.
If you pass our screening, you’ll be asked to complete one or more tests. We set the bar high and won’t extend an offer until we’re confident we’ve found the right candidate.
Thank you for considering joining DIGIPOL. If you’re excited about pushing the boundaries of AI and improving human-computer interaction, please apply now.

