EPAM Systems, Inc.
Aveiro / Global
Senior ML / Evaluation Engineer
- €70.000 - €110.000
- Remote
Aveiro / Global
We're looking for a Senior ML / Evaluation Engineer to join our team in Portugal in a fully remote working mode. In this role, you will own the design and implementation of advanced evaluation frameworks for an Enterprise Agent Development Platform—a production-grade, cloud-native ecosystem enabling scalable, secure AI agent deployment. You will create evaluation strategies that combine LLM-as-judge grading with deterministic checks, define enterprise evaluation standards, and implement CI/CD deployment gates to enforce quality metrics prior to release. This position requires strong expertise in ML system testing, evaluation design, and integration into automated pipelines for agentic environments.ResponsibilitiesDesign and implement multi-layer evaluation frameworks for agentic workflows and AI-driven applicationsBuild LLM-as-judge evaluators leveraging AWS AgentCore built-in modules and custom logic for correctness and helpfulness checksDevelop deterministic evaluators as AWS Lambda functions for rule-based validationDefine enterprise evaluation standards, including mandatory dimensions, scoring criteria, and pass/fail thresholdsImplement CI/CD deployment gates using on-demand evaluation modes to enforce quality in automated pipelinesEnable online evaluation in production by integrating sampling-based evaluation strategies and PII detection guardrailsIncorporate observability signals (OpenTelemetry spans) from AWS AgentCore into grading frameworks for trace-level assessmentGenerate metrics, logs, and dashboards from evaluation outcomes via CloudWatch or equivalent monitoring platformsCollaborate with platform, orchestration, and DevOps teams to maintain evaluation reliability and scalabilityRequirements5+ years of experience in ML engineering, AI evaluation frameworks, or AI platform developmentHands-on expertise designing LLM evaluation frameworks (LLM-as-judge and deterministic graders)Practical experience implementing CI/CD deployment gates for ML model or AI agent quality assuranceProficiency in Python for building evaluation logic (deterministic Lambda-based evaluators)Strong understanding of advanced validation dimensions, including multi-turn context integrity and workflow-level scoringNice to haveFamiliarity with AWS AgentCore Evaluations API (CreateEvaluation, GetEvaluationResult)Exposure to AWS Bedrock Guardrails for compliance and sensitive data validationExperience integrating evaluation metrics into AWS CloudWatch for monitoring and alertingKnowledge of OTel instrumentation and trace ingestion for quality scoring inputsWe offerCompetitive compensation depending on experience and skillsVariety of projects within one companyBeing a part of a project following engineering excellence standardsIndividual career path and professional growth opportunitiesInternal events and communitiesFlexible work hours
#J-18808-Ljbffr
Oliveira De Azeméis / Global
Oliveira De Azeméis / Global
Aveiro / Global
Aveiro / Global
Aveiro / Global
Aveiro / Global