- Work model
- Office
- Experience
- Not specified
- Employment
- Full Time
- Compensation
- Not disclosed
- Technology signal
- 15 tags
Technology context
15Parsed from the vacancy text; ordered by relevance to this role.
Full listing
Role description
Svitla Systems Inc. is looking for an AI QA Engineer - On-Device & Mobile Testing for a full-time position (40 hours per week) in London . Our client is an American mobile computing company. It's a central innovation team with engineers and researchers in London, Europe, the US, India, and Shanghai. They provide AI innovation and acceleration for the Business Units and are working on CV and GenAI solutions with an emphasis on device inference (Mostly Qualcomm chipsets). Researchers create/fine-tune models and hand them over to the engineering team to optimize, quantize, etc. The team then builds an application solution with customers to show value.
The position is based in London and requires visiting the office approximately three times per week.
The AI QA Engineer is a critical role responsible for ensuring the performance, safety, and reliability of our cutting-edge AI/ML models. You will be at the forefront of our development lifecycle, designing and executing comprehensive evaluation strategies to identify model weaknesses, potential biases, and critical edge cases.
This role requires a blend of analytical rigor, technical aptitude, and a deep curiosity about how AI models behave in real-world scenarios. You will not just find bugs but provide the actionable insights that drive model improvement and guide our research and development efforts.
- Proven experience in quality assurance, preferably within the AI/ML domain.
- A deep understanding of the machine learning lifecycle and the common failure modes of AI models.
- A proactive approach to collecting and curating test data to cover edge cases, with experience in data validation and managing large datasets.
- Hands-on experience with any modern UI automation framework (e.g., Playwright, Cypress, Robot Framework, Selenium, WebdriverIO).
- Experience with mobile test automation using Appium or a similar mobile automation tool.
- API testing experience (e.g., Postman).
- Experience with scripting and test automation (Python preferred, but other languages such as Java, JavaScript, or Ruby are also suitable). Ability to design, execute, and maintain automated test scenarios and validation workflows.
- Hands-on experience testing on physical Android devices: installing APKs, updating software, and performing both manual and automated testing on real devices. Experience with emulators or cloud devices alone is not sufficient.
- Strong analytical and problem-solving skills, with meticulous attention to detail and the ability to identify patterns in complex data.
- Experience with bug tracking systems (e.g., Jira) and test case management tools.
- Familiarity with using AI tools to create agents and skills and to generate code that works for testing on-device apps. Preferred tool: Claude Code.
- Familiarity with computer vision or other specific AI domains relevant to our work.
- Native Android tooling (Android Studio, ADB).
- JMeter or other performance testing experience.
Evaluation Strategy & Benchmark Development
- Design, develop, and maintain a comprehensive suite of test cases and evaluation benchmarks. Proactively identify potential model failure points, including edge cases, adversarial inputs, and sources of bias.
Error Analysis & Failure Triage
- Conduct systematic error analysis to categorize model failures and identify underlying patterns. Triage defects, prioritize them based on severity and impact, and work with the development team to ensure resolution.
Data Sourcing & Curation
- Source, curate, and manage high-quality datasets for model evaluation and testing. This includes performing data annotation and validation to ensure the integrity of our ground-truth data.
Exploratory & Adversarial Testing (Red Teaming)
- Perform unscripted, exploratory testing to discover unexpected model behaviors. Participate in red teaming exercises to intentionally challenge our models and identify potential safety and security vulnerabilities.
Functional Testing & Automation
- Validate AI-powered applications, APIs, and end-to-end user workflows to ensure functionality, usability, and reliability. Design, implement, and maintain automated tests for application-level validation using modern automation frameworks and AI-assisted tooling where appropriate.
Test Environment Management
- Set up, maintain, and troubleshoot testing and demonstration environments to ensure a stable and reliable evaluation pipeline.
Reporting & Insights
- Analyze and synthesize test results into clear, actionable reports for both technical and non-technical stakeholders. Translate complex findings into concrete recommendations for model improvement.
Process Improvement
- Actively participate in post-hoc evaluation reviews and contribute to the continuous improvement of our testing methodologies, tools, and overall quality assurance processes.