Senior SDET · AI Quality Engineer

Hey! I am Rana Usman Shahid, Senior SDET & AI Quality Engineer

Your AI fails silently. Wrong answers, hallucinated facts, broken context that your error logs never show. For 6+ years I've been building quality systems for production AI, covering LLM evaluation, RAG testing, agent quality, and performance under load, alongside high-stakes testing in FinTech and cybersecurity where silent failures cost real money.

0 Years Experience
0 Projects Delivered
0 Prompts Evaluated
About Me

Building quality systems for production AI, where silent failures cost millions.

Quality Assurance

I'm a Senior SDET with 6+ years working across production AI, covering LLM evaluation, RAG testing, agent quality, and prompt regression, as well as high-stakes testing in FinTech and cybersecurity. I've tested institutional investment management systems handling $100M+ in assets, AI governance platforms (guardrails, PII filtering, observability, access controls), conversational AI agents, and semantic search that hit 94% relevance across 10,000+ queries. My main tools are Playwright, Appium, K6, Postman, REST Assured, JavaScript, Python, and AWS CloudWatch.

I've reviewed 1,200+ prompts, found failure patterns the ML teams didn't know existed, pushed chatbot accuracy from 78% to 92%, and cut critical production defects by 85%. The work isn't really about writing more tests. It's about asking the uncomfortable questions before something ships, not after.

0 Fewer critical production defects
00 Chatbot accuracy lift
0 Semantic search relevance
$0 AUM systems tested
Skills

My Core Competencies

AI Quality & Evaluation

LLM Evaluation RAG Testing Agent Quality Prompt Regression Semantic Search Testing Chatbot Testing Hallucination Identification Guardrails & PII Filtering Bias Detection

Automation & Frameworks

Playwright Appium K6 REST Assured API Testing

Testing Types

Load & Performance Manual Testing Regression Exploratory & Risk-Based Functional Cross-browser Usability

Languages & Tools

JavaScript TypeScript Python SQL Postman AWS CloudWatch JIRA

Methodologies

Agile/Scrum TDD BDD Shift-Left Testing

Documentation

Test Plans Test Cases Bug Reports Test Scenarios

Platforms

Web App Mobile (iOS & Android)

OS / Browsers

Windows macOS Chrome Firefox Safari Edge

Let's Work Together

Have an AI system that needs to be trusted before it ships? Let's find what breaks before your users do.

Contact Me At: +92 311 4502708
Contact Me
Experience

Where I've Worked

April 2026 to Present

Senior Software Development Engineer in Test (SDET)

Kualimate

  • Working at Kualimate, an AI quality engineering practice focused on helping teams ship better AI. We sit between the model and the end user, catching the problems that unit tests simply can't find. Work covers LLM evaluation, RAG testing, agent quality, prompt regression, and pre-launch AI audits for both B2B and B2C clients.
  • Running the load and performance testing program for a live LLM-powered virtual agent. That means simulating real concurrent conversations, finding where latency starts to degrade, and giving the team the numbers they need for capacity planning.
  • Built a K6-based load testing framework in JavaScript tailored for AI workloads. It tracks metrics like Time To First Token, error rate under sustained traffic, and tail latency, which standard frameworks just don't give you out of the box.
  • Putting together a Python evaluation pipeline that scores agent responses across five areas: accuracy, relevance, safety, instruction following, and conversational quality. The goal is to catch regressions before they reach users.
  • Clients come from a mix of industries including enterprise AI, FinTech, and SaaS on the B2B side, and consumer AI products on the B2C side, spread across North America, Europe, and Australia.
January 2025 to Present

Senior Software Quality Assurance Engineer

CodingCops

  • Leading QA across four AI and consumer products at the same time, managing a team of 12 QA engineers and pushing for shift-left and risk-based testing to become part of how the team actually delivers.
  • Built validation frameworks for semantic search and chatbot systems, tested over 1,200 prompts, and shared 80+ hallucination patterns with the ML teams so they could act on them.
  • Pushed chatbot accuracy from 78% to 92% through structured prompt regression testing, and got semantic search relevance up to 94% across more than 10,000 queries.
  • Tested chatbot access controls under RBAC and ABAC, which led to uncovering 120+ critical defects including 15 high-severity security issues that needed fixing before launch.
  • Set up Playwright and Appium automation that cut regression cycles from 12 hours down to 3. That let the team ship twice a week while keeping a 98% pre-release defect resolution rate.
November 2022 to January 2025

Software Quality Assurance Engineer

Techverx

  • Was the QA owner for four enterprise platforms in heavily regulated environments. The biggest was a $100M+ AUM investment management system that went 18 months in a row without a single critical defect.
  • Caught 5 critical financial calculation bugs in portfolio valuation, trade execution, and interest accrual before they ever made it to production.
  • Ran 1,200+ test cases every quarter with a 96% pass rate, using risk-based prioritization to stay on top of multiple concurrent streams.
  • Tested 30+ payment and financial API endpoints in Postman, covering data integrity, authentication, error handling, and performance.
January 2021 to October 2022

Software Quality Assurance Engineer

LeapSofts

  • Handled QA for 8 high-traffic eCommerce platforms doing $2M+ in monthly revenue. My focus was checkout flows, inventory systems, and payment gateway integrations.
  • Caught more than 25 critical payment integration bugs before launch, protecting an estimated $200K+ in transaction revenue from going sideways.
  • Kept consistent performance for 95% of active users by testing across 15+ browser and device combinations. The usability findings that came out of it also contributed to a 20% lift in customer satisfaction.
Portfolio

My Recent Projects

Testimonials

Trusted By Multiple Clients

Want to Hire Me?

Let's Work Together On a Project

Tell me about your AI system or product. I'll tell you where it's most likely to fail and how we can prove it won't.