Senior QA · Poland · Nine hours ahead of Pacific

We're the machine it doesn't work on.

It ran once, for one person — that's not the same as working. We're a senior QA team in Poland, named after the last thing anyone says before an incident. We get U.S. startups through the three things that stall enterprise deals: security questionnaires, accessibility demand letters, and AI features nobody knows how to test.

Audits from $2,500 — fixed price, no discovery call required to get a number.
Guaranteed 4-hour daily overlap with Pacific time. Async by default.
Invoiced in USD, paid by wire or Wise. MSA under Delaware law if your counsel prefers it.
release reportrun 0417 · pre-launch
$ womm audit --target your-app.com --scope critical-flows
Checkout — happy path, 4 card typesStripe test mode · desktop + iOS SafariPass
Login, SSO and password resetGoogle + Okta · session expiry verifiedPass
Cart is unreachable by keyboardWCAG 2.2 AA · 2.1.1 Keyboard · blocks screen-reader checkoutCritical
No automated regression on release0 of 14 revenue-critical flows coveredHigh
Support assistant leaks its system promptOWASP LLM01 · reproduced on 3 of 20 adversarial inputsHigh
!CI suite is flaky18% flake rate · engineers have started re-running until greenMedium
6 checks1 critical3 high1 mediumreport delivered 08:41 PT

A sample of the first report we hand back. Yours arrives with your bugs, your severity counts, and reproduction steps an engineer can act on the same morning.

3,117

Federal website-accessibility lawsuits filed in the U.S. in 2025 — up 27% year over year. E-commerce is roughly 70% of them.

Seyfarth Shaw, ADA Title III report
86%

Of AI-generated code samples tested for cross-site scripting came back insecure. Log injection: 88%. Bigger models did not do better.

Veracode, 2025 GenAI Code Security Report
$25–75k

Typical cost of settling a single accessibility claim — before you pay anyone to actually fix the site.

Reported settlement ranges, 2025

What we're hired for

Three problems that show up right after a good quarter

Nobody buys "testing." Founders buy their way out of a specific blocker with a date attached. These are the three we're built for — and the severity tag is how we'd rate the risk of ignoring them.

Critical · blocks revenue

The security questionnaire

Your first real enterprise buyer sends a 200-line questionnaire and asks for SOC 2. Suddenly a pentest and a documented QA process are between you and the contract.

  • Web + API penetration test, scoped for your stage
  • Regression coverage on the flows auditors ask about
  • QA process documented the way an auditor reads it
  • Evidence pack that plugs into Vanta or Drata
High · legally timed

The demand letter

A letter arrives citing the ADA, or an RFP asks for a VPAT you don't have. The European Accessibility Act has applied to non-EU sellers since June 2025, so this now comes from both directions.

  • Manual WCAG 2.2 AA audit with assistive tech, not just a scanner
  • VPAT / ACR written to survive review
  • Remediation guidance your devs can ticket directly
  • Retest and a signed-off statement of conformance

How the accessibility audit works →

High · reputational

The feature that hallucinates

You shipped an LLM feature in three weeks. Now it leaks the system prompt on a crafted input, invents refund policy, and there's no way to tell if last week's prompt change made it worse.

  • Eval suites that turn "seems fine" into a pass/fail gate
  • Red teaming against the OWASP Top 10 for LLM Applications
  • RAG grounding and hallucination checks on your own data
  • Regression runs on every prompt and model change

Start here

Fixed price, fixed scope, no rate card to negotiate

Every engagement starts with a paid audit. You get a real deliverable for a known number, and you find out how we work before anyone signs a retainer.

QA Health Check

$2,500 – $4,000

We test your app and your process, then hand back a prioritized report: what's broken, what's likely to break, and what to fix first.

2 weeks

Accessibility Audit + VPAT

$3,000 – $5,000

Manual WCAG 2.2 AA audit of up to 15 screens with screen readers and keyboard-only passes, plus a VPAT you can send to procurement.

Process and deliverables →

2–3 weeks

AI Eval Sprint

$4,000 – $6,000

We build eval suites for your LLM features and run first-round red teaming, so prompt changes stop being a leap of faith.

2–3 weeks

Release-Ready Sprint

$3,500

One week of exploratory and regression testing ahead of a launch, with daily findings instead of a report at the end.

1 week

Starter

$3,500/ month

  • One dedicated QA engineer
  • Regression and exploratory passes each sprint
  • Weekly written report
  • Bugs filed straight into your tracker

Growth · most common

$6,500/ month

  • A pair: manual QA plus automation engineer
  • Playwright suite in your CI, code stays yours
  • Live dashboard and agreed SLAs
  • Overnight triage on the flows that matter

Scale

$9,500/ month

  • Managed squad: QA lead plus two engineers
  • Performance, security and accessibility in scope
  • Priority support window
  • Quarterly quality review with your CTO
Our guarantee

80% automated coverage of your critical flows within 90 days, or we keep working at no charge until you have it. If a Health Check turns up nothing worth fixing, you don't pay for it.

Need people rather than a service? Embedded engineers run $6,000–$12,000 per month depending on seniority, from manual QA through SDET.

How it works

Small commitment first, then as much as you need

The order matters. Each step earns the next one, and you can stop after any of them.

01 — Audit

Two weeks, one price

We scope over a 15-minute call, sign an NDA and MSA, and start. You get a report with severity ratings, reproduction steps and a 30/60/90 remediation plan. Many teams take that plan and run it themselves.

02 — Retainer

Quality that doesn't reset each sprint

The same engineers stay on your product. Regression suites get built once and maintained, releases get a gate, and you get a report every week instead of a surprise every quarter.

03 — Squad

More hands, same standard

When the retainer stops being enough, we embed engineers directly in your team — your standups, your board, your definition of done. We hire and train them here so you don't have to.

The time zone question

Nine hours behind you is a feature, if it's scheduled

Warsaw is nine hours ahead of Pacific time. That means your evening merge gets tested while you sleep — and it means we schedule a real overlap window rather than pretending distance doesn't exist.

Warsaw
09:00 — 19:00 CET
Overlap
17:00–21:00 CET = 08:00–12:00 PT
San Francisco
08:00 — 15:00 PT →
00:00 CET06:0012:0018:0024:00

Seniority, not headcount

Poland has one of the deepest benches of automation and SDET talent outside the U.S. Everyone on your account is senior enough to argue with your engineers, in English, about whether a bug is really a bug.

Priced between, deliberately

Well under a U.S. boutique at $60–150 an hour, and well above the $15–30 offshore tier. We're not the cheap option and we don't pitch like one.

GDPR is the default, not an upgrade

EU jurisdiction, a DPA signed as standard, and IP protection under EU law. If you sell into Europe, your QA vendor is already compliant with the regime your customers care about.

Async by design

Written findings, Loom walkthroughs, a live dashboard and a fixed overlap window. You should never have to wait for a call to know what's broken.

What you actually receive

The report, in order

Written so a founder can read page one and an engineer can work from page four.

  1. Executive summary — one page, written for the person who signs
  2. Scope and method, including what we deliberately didn't test
  3. Findings by severity, each with repro steps, evidence and a fix
  4. Metrics: coverage, defect density, flake rate, mean time to repair
  5. Process review — CI/CD, environments, test data
  6. A 30/60/90 remediation roadmap with owners
  7. For accessibility work: WCAG 2.2 AA mapping and a completed VPAT

Who's behind it

People who speak at the conferences your team reads

Not a sales team with engineers somewhere behind it. The person on your first call is the person reviewing your report.

What we're working on

We're building a public dataset on defect rates in AI-generated code — what Cursor and Copilot actually ship, measured across real repositories. First findings go out in our newsletter.

Talk recordings and slides on request

Open source

Whatever we build lives inside your repository — Playwright, Appium, your CI. There's no platform of ours to keep paying for and nothing to migrate off if you stop working with us.

No vendor lock-in, ever

Straight answers

Ask us anything on the first call — scope, method, who exactly would work on your product, what we'd refuse to do. As soon as our first U.S. engagement wraps, we'll publish it as a case study with the client's numbers, not ours.

Ask on the first call

Before you ask

The questions every U.S. founder asks us

Is the name a joke?

It's a quote. "Works on my machine" is what an engineer says right before a release turns into an incident — the moment where someone confuses it ran once, for me with it works. That gap is the entire job. We're named after it because we'd rather be the ones who close it than the ones who say it.

Who am I actually paying, and how?

Us, directly. We're an EU company and invoice in USD — payment by wire or Wise, typically settled in one to three business days. Contracts are an MSA plus a per-project SOW, with a mutual NDA and a GDPR data processing addendum, and we'll sign under Delaware law if your legal team prefers it. You'll get a completed W-8BEN-E for your records; because the work is performed outside the U.S., there's normally no withholding to handle on your side.

Nine hours is a lot. How do we actually work together?

Two ways. First, a guaranteed overlap window of 08:00–12:00 Pacific every working day for standups, pairing and anything urgent. Second, everything outside that window is async: findings land in Slack and your tracker as we make them, with a Loom walkthrough when something needs showing rather than telling. In practice most teams find the gap useful — you merge in the evening and the results are there before your first meeting.

Do you need access to production data?

Usually not. We work from staging with synthetic or anonymized data by default. Where production access is unavoidable, we scope it narrowly, sign a DPA, and can work under HIPAA or SOC 2 controls if your program requires it.

What happens to the tests you write?

They live in your repository, in Playwright or Appium, running in your CI. There's no platform of ours to subscribe to and nothing to migrate off if you stop working with us. That's deliberate — it keeps us accountable for the work rather than the lock-in.

Can you scale up if we grow fast?

Yes. Embedded engineers are our third tier, and we recruit from a deep senior pool in Poland. Realistically, adding a senior engineer to an existing account takes two to four weeks, because we won't put someone unvetted on your product to hit a date.

Isn't AI going to make QA unnecessary?

The opposite is happening. Teams using Cursor and Copilot ship far more code per week, and the defect rate per unit of code has gone up rather than down — Veracode found AI models produced insecure code for cross-site scripting in 86% of tested cases. More code, written faster, reviewed less. We use AI heavily in delivery, for test generation, triage and self-healing suites. What we don't do is pretend it removes the need for someone to exercise judgment about your product.

Next step

Fifteen minutes, and you'll know where you stand

Tell us what's coming — a funding round, an enterprise deal, a launch, a letter — and we'll tell you what we'd test first and what it would cost. If we're not the right fit, we'll say so on the call.

Book a 15-min gap call

hello@worksonmymachine.com
Replies within one business day, always from an engineer.
Free with the call: a five-issue accessibility scan of your homepage.