What one candidate assessment revealed about accessibility, UX friction, and measuring the right skills.
Recently, I completed a hiring assessment for a QA role. By the end of it, I found myself doing QA on the assessment itself.
The assessment was meant to evaluate whether I could test software, communicate clearly, and solve problems. Some parts of it did capture useful signal. I was pleased to have scored in the 97th percentile for communication because clear communication is something I care about deeply. Not just in QA work, but in the way people understand systems, move through them, and work together. It was also the section where the assessment felt most aligned with the skill being measured. I had enough room to understand the context and answer thoughtfully, instead of having to calculate, diagram, or decode several layers of the question before the timer moved on.
But other parts of the experience raised real questions for me about assessment design, accessibility, and whether certain hiring tools measure the skills they claim to measure.
The Problem Solving section was the clearest example. It included multi-step logic, scheduling, ordering, and algebra-style questions under a strict time limit. Some questions required reading carefully, translating the wording, setting up the logic, calculating or arranging information, and choosing an answer, all in under a minute.
I understand the purpose of a general logic section. Problem-solving matters in technical work, and it makes sense to evaluate reasoning outside of one narrow job function. The issue was not that the questions were general. The issue was that the format made it difficult to tell whether the test was measuring reasoning ability or rapid prompt decoding under pressure.
That distinction matters.
Good problem-solving, in QA or otherwise, depends on understanding the problem clearly enough to reason through it. In QA work specifically, that often means investigating ambiguous behavior, isolating variables, reproducing issues, clarifying requirements, communicating risk, and helping the team understand what is actually happening.
In real work, pressure does not remove the need for clarity. When a critical issue appears close to release, the goal is not to guess quickly and move on. It is to slow the problem down enough to understand it, document the risk, and help the team make a clear decision.
If a candidate has to spend most of the allotted time decoding the question itself, the assessment signal becomes harder to trust. At that point, it may be measuring speed, working-memory strain, and test-format familiarity more than practical reasoning.
The accessibility piece stood out as well. The honesty agreement restricted outside tools and browser extensions, and the guidance around screen readers was not clear enough for me to feel comfortable using one. I prefer audio support when reading dense prompts, but I avoided using a screen reader because I did not want to risk violating the rules.
At the same time, the assessment included questions that would have benefited from basic tools like a calculator, scratchpad, or movable interface.
That creates a strange contradiction: candidates are told not to use outside tools, but the platform does not provide enough built-in support for the kinds of tasks being asked.
From a product and accessibility standpoint, a few changes would have improved the experience:
So after the assessment, I did the most QA thing I could think to do: I wrote up the friction points and tried to get the feedback routed somewhere useful.
That became its own issue.
I tried to submit additional feedback to the assessment platform. At first, the support bot kept directing me back to the feedback option that appears during or immediately after the assessment. The problem was that I had already completed the assessment and could no longer access that form.
That became part of the candidate experience concern.
If candidates can only provide feedback while they are still inside the assessment flow, or immediately after a timed and stressful testing experience, they may not have enough distance from the experience to explain what happened clearly. Some usability and accessibility concerns become easier to name after the pressure has passed.
So I kept clarifying the request through the only help path I had access to: their AI support bot.
I explained that the feedback was about the assessment platform, not the hiring company. I clarified that it was candidate experience and accessibility feedback, not an application-status question. I asked whether the chat transcript could be retained or routed to the appropriate team. When the bot said there was no route, I named that limitation itself as part of the issue.
Eventually, by continuing to clarify the request and selecting that I still needed help, the bot escalated the conversation and created a support ticket.
Once the issue was successfully escalated to a support ticket, the human response was thoughtful. They thanked me for the feedback, acknowledged the accessibility and candidate experience concerns, and said they would share it with the appropriate team for review. I appreciated that response. The problem was not that no one cared once the feedback reached a person. The problem was how hard it was to find that path in the first place.
That mattered to me because it showed the difference between frustration and communication. The process was frustrating, but I was not trying to argue with a bot. I was trying to get a product issue routed correctly.
The process became its own little QA exercise:
And suddenly, there it was. I had just completed an assessment for a QA role, and then my first real QA-style ticket in months was about the assessment platform itself.
There was something admittedly satisfying about turning the whole experience into a case study of its own: looking at it the way I would look at any other system and asking where the user path broke, what made the task harder than it needed to be, and how it could be improved.
Hiring assessments are not neutral objects. They are designed experiences that candidates have to move through: instructions, constraints, timing, interface, and feedback paths. When those pieces introduce unnecessary friction, the assessment may start measuring the experience around the skill as much as the skill itself.
A good assessment should measure the intended skill as cleanly as possible. It can be challenging without introducing unnecessary friction. The goal should be to evaluate the candidate’s ability, not their ability to work around unclear wording, missing tools, or avoidable stress.
For an assessment platform to evaluate people well, it must also be willing to evaluate its own candidate experience and make space for that feedback with the same care.
That is the kind of QA work I care about: not only finding defects, but understanding where a system creates confusion, where users lose trust, and where the experience stops supporting the outcome it was designed to produce.