Senior Software Engineer – Python (LLM Evaluation & Repository Validation)
Senior Software Engineer – Python (LLM Evaluation & Repository Validation)
- Location
- Remote — Worldwide
- Engagement
- Contractor
Earns 25 points on this device — once per role per day
Applications are handled by Turing on their own site. Dealuxe is not the employer and does not screen applicants.
Shape the cognitive capabilities of next-generation artificial intelligence models by engineering verifiable software evaluation datasets with Turing.
The evolution of artificial intelligence has reached a critical juncture. Large Language Models (LLMs) can generate fluent text, but ensuring they can execute complex, real-world software engineering tasks remains one of the industry's greatest frontiers. Grounded software validation requires rigorous human expertise to build reliable benchmarks that test how AI interacts with actual source code, manages dependencies, and solves authentic coding bugs.
Turing—one of the world's fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems—is engaging senior Python software engineers for a high-impact role: Senior Software Engineer – Python (LLM Evaluation & Repository Validation). This remote contractor opportunity invites you to combine elite software engineering with cutting-edge AI research.
About the Project and Role
In this project, engineers build verifiable software engineering (SWE) tasks based on public repository histories using a synthetic approach with human-in-the-loop oversight. The objective is to expand dataset coverage across diverse programming languages, difficulty tiers, and architectural domains.
As a tech-lead-level contributor, your role involves hands-on software engineering work including development environment automation, issue triaging, and evaluating test coverage and quality. You will define how frontier AI models reason through complex codebases.
What Your Day-to-Day Looks Like
Your daily responsibilities blend deep technical engineering with analytical evaluation:
- Issue Triaging: Analyze and triage GitHub issues across trending open-source libraries.
- Environment Automation: Set up and configure code repositories, including Dockerization and local environment setup.
- Quality Assurance: Evaluate unit test coverage and code execution quality.
- Local Testing: Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.
- Research Collaboration: Work alongside AI researchers to design and identify repositories and issues that present genuine challenges for LLMs.
- Leadership: Opportunities to lead a team of junior engineers collaborating on shared project milestones.
Are You a Fit? Required Skills & Experience
To succeed in this rigorous engagement, candidates should meet the following technical benchmarks:
- Experience: Minimum 3+ years of overall professional software engineering experience.
- Language Proficiency: Strong, demonstrated experience with **Python**.
- Tooling Mastery: Proficiency with Git, Docker, and basic software pipeline setup.
- Codebase Navigation: Exceptional ability to understand, navigate, and debug complex codebases.
- Local Execution: Comfortable running, modifying, and testing real-world projects locally.
- Open Source: Previous experience contributing to or evaluating open-source projects is a strong plus.
- Nice-to-Haves: Participation in LLM research/evaluation or experience building developer tools and automation agents.
Engagement Details & Location Eligibility
This is a 3-month remote contract assignment with an expected start date next week. Candidates can choose among three time-commitment tiers: 20 hours/week, 30 hours/week, or 40 hours/week (minimum 4 hours per day, with at least a 4-hour daily overlap with PST).
Eligible Regions: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.
The Evaluation Process
To ensure high standards of technical excellence, the evaluation process is streamlined yet thorough (approximately 75 minutes total):
- Technical Round (60 minutes): In-depth technical assessment covering Python, system design, and debugging.
- Technical & Cultural Discussion (30 minutes): Collaboration, communication, and cultural alignment review.
Maximizing Your Success: Onboarding & Stability
Once selected, completing your onboarding promptly and diving into your first 10 hours of work is essential. Fulfilling this initial milestone guarantees project stability, solidifies your standing within the network, and unlocks access to more advanced tasks and ongoing assignments.
Ready to Shape the Future of AI?
Join Turing's elite network of software engineers and evaluate next-generation AI systems.
Submit your application through the secure official portal below.
What the work is
- Issue Triaging: Analyze and triage GitHub issues across trending open-source libraries.
- Environment Automation: Set up and configure code repositories, including Dockerization and local environment setup.
- Quality Assurance: Evaluate unit test coverage and code execution quality.
- Local Testing: Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.
- Research Collaboration: Work alongside AI researchers to design and identify repositories and issues that present genuine challenges for LLMs.
- Leadership: Opportunities to lead a team of junior engineers collaborating on shared project milestones.
Ready to apply for Senior Software Engineer – Python (LLM Evaluation & Repository Validation)?
The application is on Turing's own site and takes a few minutes.
Earns 25 points on this device — once per role per day
Dealuxe is not the employer, does not set the pay or the hiring terms, and cannot guarantee a role is still open. If you complete a purchase or form, we may earn a small commission at no extra cost to you.
Following an offer here banks 10 points on this device — once per page, within the 500 points a day anything on the site can earn.Ad Disclosure: the application link is a referral link.
While you job-hunt, save on the brands you already use
Browse 2,287 vetted brands with live commission offers and exclusive deals — every one open to everyone, no sign-in needed.
Explore all brand categories →Scroll through the piece and stay a moment. Reading pays 5 points and sharing pays 50. Following the apply link pays 25, and buying coins pays back 25 points a dollar.
Get stories in your inbox
New brand drops, deal breakdowns and the best of the Journal, straight from The Storefront Blog. Free forever — and subscribing pays you 25 points.
Drop your email in the box above, then bank the bonus. A paid plan pays 75 — three times the free tier — plus every paid-only post.
Copies the link with your caption. Grab a username to bank points across devices.
Boost this listing
See what's trendingTrade the points you've earned to push this up Trending and the homepage, where more readers will find it.
Similar roles
Docker Data Validation Engineer
Turing
An in-depth look at Turing's mission, day-to-day responsibilities, technical requirements, engagement logistics, and onboarding milestones for containerization engineers.
- Location
- Remote — Global
- Posted
Engineering Manager & Delivery Leader
Turing
Lead large-scale technical teams, drive Supervised Fine-Tuning (SFT) and RLHF data pipelines, and bridge elite engineering with cutting-edge artificial intelligence research.
- Location
- Remote — LATAM
- Posted
SciCode Trainer
Turing
Engage in high-impact remote contracting by authoring complex mathematical datasets to train next-generation artificial intelligence models.
- Location
- Remote — Global
- Posted
Bridging Materials Science and Artificial Intelligence: The Turing SciCode Masterclass
Turing
Explore how elite domain experts are shaping frontier AI models through rigorous scientific coding benchmarks, advanced Python simulations, and structured problem architecture.
- Location
- Remote — Global
- Posted
Scientific Coding - Physics and Python: Shaping Frontier AI Benchmarks
Turing
Leverage your advanced physics expertise and programming mastery to train next-generation artificial intelligence models through Turing's rigorous SciCode initiative.
- Location
- Remote — Global
- Posted
Data Scientist / Analyst
Turing
Discover how senior data professionals can leverage Python and advanced analytics to train frontier models, partner with leading AI labs, and shape autonomous systems.
- Location
- Remote — Global
- Posted
More software & data listings
- MLE Bench – Data Analyst
- SWE Bench Data Engineer & Data Scientist
- Python Machine Learning Engineer
- LLM Go Developer
- Senior Python Developer
- JavaScript / TypeScript Full-Stack Developer
- Building the Next Generation of Dialog Agents
- LLM C/C++ Developer
Other roles at Turing
- AI Quality Analyst (Personalization) für den deutschen Markt
- Technical Content Writer
- Peluang Karier Global: Menjadi Business Analyst Bahasa Indonesia di Turing untuk Mengembangkan AI Masa Depan
- Gabay sa Pagpasok bilang Business Analyst (Tagalog Language) sa Turing
- Music and Audio Expert
- Illustrator, Sketcher & Cartoonist
Be the first to comment
Loading comments…