# Neuroscale — full content > Neuroscale builds Arbi, the AI recruiting platform that transforms talent acquisition into a science. The index of this content, with page links, lives at https://neuroscale.ai/llms.txt. --- # Blog posts # The screening bottleneck is a measurement problem > Recruiters are not slow readers. They are being asked to hold a consistent standard across four hundred profiles with nothing to hold it in. Here is what changes when the standard becomes something the system can apply. - Author: Dana Whitfield, Head of Product at Neuroscale - Published: 2026-08-05 - Category: Screening - Canonical: https://neuroscale.ai/blog/screening-is-a-measurement-problem Ask a recruiter why screening takes so long and you will hear about volume. Four hundred applicants, one req, one person. The number is real, but it is not the reason. Reading four hundred profiles at ninety seconds each is nine hours of work: a long day, not an impossible one. The reason screening takes so long is that nobody can tell you when it is finished. There is no line the pile is measured against, so the work has no end state, only a point of exhaustion. What gets called a throughput problem is almost always a measurement problem wearing a throughput costume. - **412** — Median profiles per open req, mid-market - **6.4s** — Median time on a first-pass resume - **31%** — Rejections a second reviewer disagreed with ## The tenth resume is not the first resume The first profile of the morning gets read. The tenth gets skimmed. The hundredth gets pattern-matched against the nine before it, which is a polite way of saying it gets compared to the wrong thing. This is not a failure of diligence. It is what attention does under load, and it has been measured in radiologists, air traffic controllers, and appellate judges before anyone got around to measuring it in recruiters. The consequence in hiring is specific and expensive: the bar drifts. Not randomly, but in the direction of whatever the reviewer has just seen. Five strong backend profiles in a row and the sixth gets judged against them rather than against the role. By the end of the stage you have a shortlist that is internally inconsistent in ways nobody can reconstruct, because the standard existed only in one person's head and it was moving the whole time. > A shortlist is a claim about a group of people. If you cannot say what the claim was measured against, you have produced a preference, not a decision. ## What a criterion has to do to be useful The instinct, once you accept that the standard needs writing down, is to write down what you already say out loud. That is where most scorecards die. "Strong engineer." "Good communicator." "Startup mindset." These read like requirements and function like mirrors. Every reviewer sees their own definition in them, so the scorecard produces the same drift it was meant to prevent, now with a paper trail. A criterion earns its place when two people reading the same profile would reach the same verdict on it. That is a high bar and it rules out most of what ends up on an intake form. | Instead of | Write | | --- | --- | | Strong backend engineer | Has owned a service in production handling meaningful traffic, not only feature work inside someone else's service | | Startup experience | Was employee number one to fifty at a company under 200 people, for at least eighteen months | | Good with data | Has written and maintained SQL against a production warehouse, not only consumed dashboards | | Leadership potential | Has been the named technical owner of a project involving at least three other engineers | The right-hand column is longer, uglier, and testable. That last property is the only one that matters. You will also notice that writing it forces an argument with the hiring manager that would otherwise have surfaced in week five, when the first shortlist gets rejected for reasons nobody articulated in the intake. > **A useful test** > > Read the criterion, then ask what a profile would have to contain for you to mark it a fail. If you cannot answer in one sentence, the criterion is a mood, and it will be applied differently every time it is used. ## Three ways a written standard still fails Writing the criteria down is necessary and not sufficient. There are three reliable ways it still falls apart. 1. **The criteria are never re-read.** They get written at intake, and by profile forty the reviewer is working from memory again. A standard that lives in a document nobody has open is a standard that is not being applied. 2. **Everything is weighted the same.** Nine criteria, all mandatory, means the ninth one (usually something like "based in a compatible time zone") knocks out candidates who are exceptional on the first three. Ordering by importance is not a nicety; without it a checklist optimises for the inoffensive. 3. **The verdicts are not recorded.** If the output is a yes or a no with no note attached, the reasoning evaporates. Six weeks later, when the hiring manager asks why a particular person was not advanced, the honest answer is that nobody knows. Each of these is a bookkeeping failure, and bookkeeping is precisely the kind of work that people are bad at and software is good at. ## What changes when the machine holds the standard Once criteria are explicit, ordered, and applied by something that does not get tired, the shape of the work changes. The reviewer stops being the instrument and becomes the person reading the instrument. Concretely: every profile in the stage is evaluated against every criterion, the results come back as a percentage plus a per-requirement verdict, and the pile arrives sorted. The nine hours of reading do not disappear. They get spent differently. Instead of ninety seconds on all four hundred, you spend fifteen minutes on the forty at the top and twenty minutes on the boundary cases in the middle, which is where the actual judgement lives. ![A candidate review drawer showing a match percentage and the evidence behind each requirement](https://neuroscale.ai/screening/review@2x.webp) *Every verdict carries the passage from the profile it was drawn from. The disagreement you want is with the evidence, not with the number.* The second change is subtler and more valuable. Because the standard is written and the verdicts are recorded, you can audit the standard itself. Sort the rejects, read twenty of them, and it becomes obvious within minutes whether a criterion is doing what you meant it to do. Ours regularly are not. The fix takes thirty seconds and re-running the stage takes one click, which is the first time in most recruiting workflows that being wrong has been cheap. ## Where this still needs a person None of the above decides anything. It cannot, and the moment a screening tool starts silently discarding profiles on your behalf you have lost the only property that made writing the criteria worthwhile: that a human can check the work. There are also things a criterion will never capture. A candidate who has done something adjacent and unusual, who is early in a steep trajectory, who wrote a cover letter that tells you more than the six roles above it. Those are found by reading, and they are found more often when the reader still has attention left at profile three hundred. That is the trade the whole thing rests on. Machines are good at applying the same standard four hundred times. People are good at noticing the profile the standard was never designed for. Screening breaks when you ask either one to do the other's job. --- # How recruiters actually read a resume > We watched forty recruiters review the same twenty profiles. The order they read in, the things they never looked at, and why the tenth resume never gets the attention the first one did. - Author: Priya Raghunathan, Research Lead at Neuroscale - Published: 2026-07-22 - Category: Research - Canonical: https://neuroscale.ai/blog/how-recruiters-actually-read-resumes We asked forty recruiters to review the same twenty profiles for the same fictional role, then watched where their eyes went and in what order. Half were agency, half in-house. Experience ranged from eighteen months to nineteen years. The point was not to catch anyone out. It was to find out what a resume review actually is, because every product decision we make about screening rests on an assumption about that, and we would rather the assumption were checked. ## Nobody reads top to bottom Not one participant read a profile in document order. The dominant pattern, in thirty-three of forty sessions, was: - Current title and current company - The list of company names, read as a column, ignoring everything between them - Total years, computed by subtracting the earliest date from today - Then, only if the first three passed, back to the top to read anything in prose The median time spent before the first accept or reject lean was **6.4 seconds**. The median time spent on a profile that got rejected was **11 seconds** in total. Profiles that got advanced received a median of 74 seconds, which is where nearly all the reading happened. This has an obvious implication that we had somehow never stated plainly: for the large majority of candidates, the review is not a review. It is a company-name lookup with a duration check attached. ## Company names are doing most of the work When we asked participants afterwards what they had based the decision on, the answers were about skills, relevance, and trajectory. The recordings showed something narrower. Recognisable company names produced a measurable dwell increase on the rest of the profile; unrecognisable ones produced a scroll. The gap between the two columns below is the finding. Nobody was lying to us; people are simply poor witnesses to their own attention. | What they said they weighed | Said it | Gaze pattern supported it | | --- | --- | --- | | Skills and tools listed | 68% | 22% | | Relevance of recent work | 55% | 40% | | Career trajectory | 43% | 31% | | Company names | 12% | 79% | The effect was strongest in the least experienced group and did not disappear in the most experienced group. It changed shape. Senior recruiters were more likely to recognise a small company as a strong signal, which is a better heuristic, but it is still a heuristic about the employer rather than about the person. > **Why this matters more than it used to** > > Company prestige was a defensible proxy when most engineers worked at a few hundred legible companies. It is a poor one now, when the interesting work is distributed across thousands of small companies nobody outside their category has heard of, and when the same logo covers both an infrastructure team and a team writing internal CRUD. ## The attention curve is steeper than anyone admits We ordered the twenty profiles randomly per participant, which let us measure attention against position rather than against the candidate. - **74s** — Median dwell, profiles 1 to 5 - **38s** — Median dwell, profiles 6 to 12 - **19s** — Median dwell, profiles 13 to 20 Attention fell by roughly three quarters across twenty profiles. Twenty. The reqs these people work on have between two hundred and a thousand. More uncomfortable: when we re-showed three profiles late in the session that had appeared early, eleven participants gave a different verdict the second time, and nine of those eleven were more negative. Nobody noticed they had already seen the profile. ## What we changed because of it Three things, all of them narrower than the study might suggest. 1. **Rank before anyone reads.** If the first five profiles get four times the attention of the last five, the only responsible thing to do is make sure the first five are the ones most likely to deserve it. This is the entire argument for scoring a stage before review, and it is a stronger argument than the time saving that usually gets quoted. 2. **Lead with evidence, not with the company.** Our review panel now opens on the criteria and the passages that satisfied them. The employment history is one scroll away rather than the first thing in the eye's path. Small change, and it moved reviewer agreement on the same profile from 61 percent to 78 percent in a follow-up round. 3. **Show the reviewer their own drift.** Recording the verdicts is not only for the audit trail. If your accept rate over the last thirty profiles has fallen off a cliff relative to the thirty before, that is worth surfacing, because the alternative is finding out at offer stage that the back half of the pile never got a fair look. ## The thing we did not change We did not build anything that hides the resume. Several people, on hearing about the study, suggested the obvious follow-up: strip the company names, strip the dates, show only evidence against criteria. We tried it. Reviewer confidence collapsed and review time went **up**, because people started hunting for the context they had been denied. Trajectory is real information. The shape of a career tells you things no individual line does, and taking it away does not remove bias so much as relocate it somewhere you can no longer see. The better answer is not less context. It is making sure the context arrives after the evidence rather than instead of it, and that the person on profile three hundred is looking at a list that was already sorted by something other than the order the applications came in. --- # Boolean search is a 1974 answer to a 2026 problem > Keyword strings assume the words on a profile are the same words in your head. They almost never are. What it takes to search for a person instead of a string. - Author: Marcus Feld, Founding Engineer at Neuroscale - Published: 2026-07-09 - Category: Sourcing - Canonical: https://neuroscale.ai/blog/boolean-search-is-a-dead-end Boolean retrieval was formalised in the early 1970s for librarians querying bibliographic databases over teletype. It was a good design. It is still the interface most recruiters use to find people, which is roughly like navigating with a sextant because it worked for the Admiralty. The problem is not that operators are bad. The problem is what a keyword string assumes: that the words in your head are the words on the profile. In sourcing, they almost never are. ## Where the string breaks Here is a real query, lightly anonymised, for a senior infrastructure role: ``` ("site reliability" OR "SRE" OR "infrastructure engineer") AND (kubernetes OR k8s) AND (terraform OR pulumi) AND NOT (junior OR intern OR "student") ``` It is a competent string. It also silently excludes: - The engineer whose title is "Platform Engineer" because that is what her company calls the team - The engineer who has run production Kubernetes for four years but wrote "container orchestration" on his profile - The engineer who has never touched Terraform because her employer standardised on CloudFormation, and who would be productive in Terraform in a week - Anyone at a company where the infrastructure team sits inside a product org and the titles reflect the product And it silently includes anyone who put Kubernetes in a skills list after a weekend tutorial, because a keyword match cannot tell the difference between having done something and having mentioned it. > **The asymmetry that hurts** > > A false positive costs you thirty seconds. You open the profile, see the mismatch, move on. A false negative costs you the candidate, permanently and invisibly. Boolean tuning is almost always aimed at the error you can see. ## The vocabulary problem is not solvable with more OR The usual response to a miss is to widen the string. Add "Platform Engineer". Add "container orchestration". Add CloudFormation. This works for the specific miss you noticed and does nothing for the class of misses it belongs to, because you are enumerating a vocabulary that has no fixed size. Job titles in particular are a moving target. We track roughly 40,000 distinct engineering titles across the profiles we index, and the long tail is not noise. It is companies naming things after their own architecture. "Developer Experience Engineer" and "Build Systems Engineer" and "Internal Tools Engineer" are frequently the same job, and no string contains all three unless someone thought of all three. - **40k+** — Distinct engineering job titles indexed - **3.1** — Median distinct titles per actual role type - **58%** — Of qualified profiles missed by a typical string That last figure comes from a small internal exercise. We took twelve strings written by experienced sourcers, ran them against a pool where we had manually labelled who was genuinely qualified, and measured what the string returned. The median string found 42 percent of the qualified pool. The sourcers, shown the misses afterwards, agreed with the label in almost every case. ## Recall you cannot see The deeper issue is epistemic. A search interface shows you what it found. It has no way of showing you what it did not, so there is no feedback signal telling you your string is too narrow. You get results, the results look reasonable, and the eighteen people you missed never enter the conversation. Recruiters compensate with volume, running six strings instead of one and sourcing from three platforms, which raises recall a little and raises effort a lot. It also means the same person surfaces four times and gets deduplicated by hand. ## What replaces the string Not natural language search as a marketing phrase. What actually has to change is the unit of matching. A string matches tokens. What you want is a system that matches a description of a person against the evidence in a profile, which requires two things a keyword index does not have: 1. **A representation that survives paraphrase.** "Ran production Kubernetes" and "operated containerised workloads at scale" need to land in the same place. This is what embeddings are genuinely good at, and it is why a semantic index finds the Platform Engineer without anyone having thought to type "Platform Engineer". 2. **A judgement step over the retrieved set.** Retrieval gets you a candidate pool that is broad and noisy. Something then has to read each profile against your actual requirements and say why it does or does not fit. Without that second pass you have replaced a precise-and-narrow tool with a fuzzy-and-wide one, which is not obviously an improvement. The combination is what makes the difference. Broad retrieval means you stop missing people for vocabulary reasons. Evidence-based judgement over that pool means the breadth does not turn into a thousand profiles to sift. ## Keep the operators None of this is an argument for taking Boolean away. There are constraints that are genuinely binary and should be expressed as such. Work authorisation, a hard location boundary, a security clearance, a licence. Handing those to a language model is worse in every respect: slower, more expensive, and less predictable than an index lookup that has been correct since 1974. The right shape is a filter for the things that are actually filters, and a description for the things that are actually descriptions. Most sourcing tools force everything into the first category, which is why sourcers spend their afternoons writing parentheses instead of talking to people. --- # Your reply rate did not fall because you sent too few emails > Outreach volume has roughly tripled in three years and replies have halved. The arithmetic of the volume trap, and the three things that still move a reply rate. - Author: Dana Whitfield, Head of Product at Neuroscale - Published: 2026-06-24 - Category: Sequencing - Canonical: https://neuroscale.ai/blog/reply-rates-and-the-volume-trap Every team we talk to has the same shape of story. Two years ago the first-touch reply rate was somewhere around 22 percent. Now it is nine. The sequences did not get worse. The volume went up, everyone's volume went up, and the inbox on the other end did the arithmetic. The standard response is to send more, which is the one response guaranteed to make the underlying condition worse. ## The arithmetic of the volume trap Consider a candidate pool of a hundred senior engineers in a specific niche, and thirty companies hiring into that niche. If each company sends one sequence of four touches per quarter, each engineer receives 120 messages a quarter. That is more than one per working day, all of them about jobs, most of them opening with a compliment about a GitHub profile. Reply rate is not a property of your email. It is a property of your email divided by everything else in the inbox that quarter. Doubling your send doubles the denominator for everyone including yourself, and the equilibrium it moves toward is one where nobody replies to anybody and the only winners are the email providers. - **3.1x** — Increase in recruiting sends per candidate, 2023 to 2026 - **−54%** — Change in median first-touch reply rate - **1.4** — Recruiting messages per working day, senior ICs We are contributing to this. So is every other tool in the category. Pretending otherwise would be strange, and the honest position is that a sequencing product should be measured on replies per thousand messages, not messages per hour. ## What actually moves a reply We looked at 2.4 million first-touch messages sent through Arbi over eighteen months and modelled reply against everything we could measure. Three things came out with an effect size worth caring about. Most of what gets written about outreach did not. ### Specificity that could not be templated Not personalisation tokens. The model does not care that you interpolated a first name, and neither does the recipient. What moves the number is a sentence that could only have been written about that person: a reference to a specific project, a talk, a design decision, an unusual path between two roles. Messages containing at least one such sentence replied at **2.7x** the rate of messages without one. The effect held after controlling for sender, seniority, and company. ### A first message that asks for one thing Messages with a single explicit ask outperformed messages with two or more by 61 percent. The common failure is a message that asks for a reply, a call, a CV, and a referral, which reads as a form rather than a conversation and gets processed accordingly. The best-performing single ask was not "are you open to a call". It was a closed question about the person's situation that could be answered in a sentence. Low cost to answer, and answering it starts a thread. ### Length, but not the way people think Short messages do better up to a point and then stop. The curve bottoms out somewhere around 60 words and rises again past 220, and the long tail is real: detailed messages about a specific technical problem the company is facing reply well. What performs badly is the middle, where 120 words of generalities are long enough to demand attention and short enough to say nothing. > **What did not show up** > > Send day, send hour, subject-line length, emoji, follow-up count past three, and whether the sender's title said "Talent" or "Recruiting". All of these have effect sizes indistinguishable from zero in our data. Most outreach advice is about these. ## Timing matters, and it is not a hack One finding did surprise us. Reply rate against the recipient's own tenure is far from flat. Messages landing between month 20 and month 34 in a role reply at roughly double the rate of messages landing in the first year. This is not a trick to schedule around. It is a reason to build a pipeline you can wait with. The candidate who says no in March because she started in January is a strong yes eighteen months later, and the only teams that capture that are the ones where "no, not now" writes a date into a system rather than closing a tab. Most sequencing tools are built for a campaign that ends. The valuable thing is the one that does not. ## The uncomfortable conclusion If specificity is the thing that works, and specificity is expensive, then the honest version of outreach at scale is not "send more, personalised automatically". It is: send fewer, to a list that has been narrowed properly, with something real in the first paragraph. That puts the weight back on the stage before outreach. A sequence sent to 400 loosely-matched people at 4 percent yields 16 replies and burns the list. The same effort spent narrowing to 80 genuinely strong matches, with a real sentence each, yields more replies, more conversations worth having, and a pool that will still take your email next year. The volume trap is not primarily a writing problem. It is a targeting problem that shows up in the writing, because you cannot say anything specific about a person you had no real reason to contact. --- # Structured interviews, without the scorecard theatre > Everyone agrees structured interviews predict performance better. Almost nobody runs them properly, because the format asks interviewers to do bookkeeping while listening. - Author: Elena Sorokina, Talent Partner in Residence at Neuroscale - Published: 2026-06-11 - Category: Interviewing - Canonical: https://neuroscale.ai/blog/structured-interviews-without-the-theatre The evidence on structured interviews has been stable for four decades and is not seriously disputed. Ask every candidate the same questions, rate against defined anchors, and you roughly double the predictive validity of the interview relative to an unstructured conversation. It is one of the few genuinely settled findings in the field. Almost nobody does it. Not because hiring managers disagree with the research. Most of them can cite it. The reason is that the format asks a human being to run a rubric and hold a conversation at the same time, and human beings are bad at that. ## What the research actually says Two things get conflated. Structure is not one dial, it is two. - **Question structure.** Every candidate gets the same questions in the same order, drawn from the requirements of the role rather than from whatever the interviewer thought of on the way in. - **Evaluation structure.** Every answer is rated against defined levels, written down before the interview, with an example of what a two looks like versus a four. The second matters more than the first, and it is the one that gets dropped. Plenty of teams have a shared question list and a scorecard with five competencies rated one to five, with nothing anywhere defining what a three means. That is question structure with the evaluation left as an exercise for the reader, and it recovers very little of the predictive gain. > **The anchor is the whole thing** > > A rating scale without behavioural anchors does not measure the candidate. It measures the interviewer's mood, calibrated against every other candidate they happen to remember. Two interviewers using the same unanchored one-to-five scale routinely differ by a full point on the same recording. ## Why it collapses in practice Watch someone run a structured interview properly and the problem is obvious within ten minutes. They are doing four jobs at once: 1. Asking the question as written, without leading 2. Listening well enough to ask a good follow-up 3. Taking notes detailed enough to justify a rating later 4. Holding the anchors in mind so the rating is against the rubric rather than against the last candidate Three and four lose. Notes degrade into fragments, ratings get filled in afterwards from memory, and the memory is dominated by the most recent five minutes and the candidate's warmth. The scorecard gets completed, the process is described as structured, and the actual mechanism that produces the validity gain never ran. We have a name for this internally. Scorecard theatre: all the artefacts of structure, none of the measurement. ## Bookkeeping is the enemy of listening The instinct is to fix this with discipline: better training, stricter templates, a reminder to fill the scorecard in within an hour. It helps at the margin and it does not survive a busy week. The realistic fix is to remove the bookkeeping from the interviewer entirely. If the interview is recorded and transcribed, then the mapping from what was said to what the rubric asks about is a retrieval problem, not a memory problem. The interviewer's only job becomes the one they are good at: asking a real question and listening to the answer. That produces a scorecard where each competency arrives with the passages that bear on it, timestamped, and the interviewer's job is to agree, disagree, or push back on evidence rather than to reconstruct an hour from four lines of handwriting. ## Making the structure invisible The design goal we ended up with is that a well-structured interview should feel less structured to the candidate, not more. This sounds contradictory and is not. What makes a structured interview feel like a deposition is the interviewer's visible bookkeeping: the eyes going to the notes, the pause while something is typed, the mechanical transition to the next item. Remove those and what is left is a conversation that happens to cover the same ground every time. The things we deliberately did **not** automate: - **Follow-ups.** The second question is where the information is, and it depends on what was just said. Scripting it defeats the purpose. - **The rating itself.** Evidence is assembled automatically. The judgement is made by the interviewer, on the record, and it is theirs. - **The decision.** A scorecard is an input to a debrief, not a replacement for one. ## Calibration is the whole game The last piece is the one teams skip and then wonder why their scores do not mean anything across interviewers. Take three recorded interviews. Have everyone who will run the loop rate them independently against the anchors. Then compare. The first time a team does this, the spread is usually two full points on at least one competency, and the conversation that follows, about what a four actually looks like on this competency for this role, is worth more than any amount of interviewer training. Do it again after twenty interviews. If the spread has not closed, the anchors are the problem, not the people. Structured interviewing is not a form to fill in. It is a shared definition of the thing being measured, maintained by argument, with the paperwork moved somewhere it cannot interrupt the listening. --- # Inside the evidence layer: how Arbi shows its work > A score nobody can audit is a rumour with a number attached. A walk through the retrieval and citation path that puts a source behind every judgement Arbi makes. - Author: Marcus Feld, Founding Engineer at Neuroscale - Published: 2026-05-28 - Category: Engineering - Canonical: https://neuroscale.ai/blog/inside-the-evidence-layer A score with nothing behind it is a rumour with a number attached. If Arbi tells you a candidate is an 84 percent match and cannot say which sentence in which document produced that, you have not saved any work. You have moved the reading from before the decision to after it, when someone asks you to justify the shortlist. This post is about the part of the system that makes the number checkable. Internally we call it the evidence layer, and it is roughly half the engineering in screening. ## The problem with a bare score The naive implementation is one prompt: here is a profile, here are the criteria, return a score. It works. It demos beautifully. It also fails in three specific ways that only show up at volume. - **It cannot be audited.** When a recruiter disagrees with a verdict, the only available response is to re-run it and hope. - **It degrades with document length.** A forty-page CV plus three years of GitHub activity does not fit comfortably in one pass, and quality falls off in the middle of long contexts in ways that are hard to detect from the output. - **It is not stable.** Two runs of the same profile against the same criteria produce different numbers, and there is no diff you can inspect to find out why. The fix for all three is the same: stop asking for a judgement about a person and start asking for a judgement about a passage. ## Retrieval before judgement Every profile is decomposed into spans: a role, a project description, a paragraph from a cover letter, a repository README, a certification record. Each one carries a stable identifier and a pointer back to its source. For each criterion, we retrieve the spans most likely to bear on it. This is hybrid: a dense vector search so that "operated containerised workloads" finds a criterion about Kubernetes, plus a lexical pass so that exact tokens like a specific licence number or framework version are not lost to paraphrase. ```text criterion → retrieve k spans (dense + lexical, reciprocal rank fusion) → judge each span independently: supports / contradicts / irrelevant → aggregate to a verdict with the supporting span ids attached ``` Judging spans independently is the important part. It bounds the context each judgement sees, it makes the unit of work small enough to run in parallel, and it means a wrong verdict can be traced to a specific span rather than to a vibe about the whole document. ## Citing a span, not a document Anyone can attach a source link. The useful thing is a character range. Every verdict carries the span ids it rested on, and every span id resolves to an offset in the original document. That is what makes the review panel work: clicking a requirement highlights the exact sentence in the CV that satisfied it, in place, with the surrounding paragraph visible. It also gives us the only regression test that matters. When a recruiter marks a verdict wrong, we capture the criterion, the spans, and the verdict as a labelled example. That corpus, currently a bit over 90,000 human-corrected judgements, is what we evaluate model and prompt changes against, and it is considerably more valuable than any public benchmark for this task. - **90k+** — Human-corrected judgements in the eval set - **6** — Median spans retrieved per criterion - **1.9%** — Verdicts overturned on recruiter review ## What we do when the evidence is thin The most consequential design decision in the whole system is what happens when retrieval comes back with nothing good. The tempting behaviour is to let the model reason from context: the candidate was at a company that certainly uses Kubernetes, in a role that would certainly involve it, so mark it a pass. This is exactly the behaviour that makes a screening tool untrustworthy, because the inference is invisible and frequently wrong. Arbi returns **not evidenced** and says so. It is a distinct state from a fail, it renders differently, and it is the correct answer surprisingly often. Resumes are lossy documents, and plenty of true things about a candidate are simply not written down anywhere in them. > **Not evidenced is a feature** > > Roughly 14 percent of criterion verdicts come back as not evidenced. Recruiters treat these as a to-do list of questions worth asking on a screen call, which turns out to be more useful than a confident guess would have been. ## Cost and latency Judging every criterion against six spans for four hundred candidates is a lot of inference. Three things keep it viable. 1. **Spans are shared across criteria.** Decomposition and embedding happen once per profile and are cached; re-running a stage with edited criteria only re-runs the judgement step. 2. **Cheap models do most of the work.** Span-level relevance is a small classification problem. It does not need a frontier model, and routing it to one is how teams end up with a screening bill larger than their ATS. 3. **Escalation is selective.** Judgements near a decision boundary, and criteria the recruiter has marked as high importance, get a second pass from a larger model. Everything else does not. Median wall-clock for a 400-profile stage against eight criteria is a little under four minutes. The recruiter who kicked it off is generally still in the tab. ## Where this is still weak Two places, both known. Cross-span reasoning is limited by design. A criterion like "has grown a team from three to fifteen" requires assembling a fact from several places in a document, and our aggregation step handles the easy version of this and misses the hard one. We would rather miss it and mark it not evidenced than hallucinate it. And the evidence layer is only as good as the source. A profile that is out of date is out of date, and no amount of retrieval fixes a document that does not mention the last two years. That is a data problem, not a modelling one, and it is the honest limit on what screening can tell you before someone picks up the phone. --- # Release notes # Arbi v2.18: Skill maps on every shortlist > A shortlist now carries a map of where its strength actually sits, so you can see at a glance that eight of your top ten are strong on the same two criteria and thin on the third. - Released: 2026-08-12 - Area: Screening - Canonical: https://neuroscale.ai/releases/2.18 ## Changes - **New**: Skill maps render for any stage with four or more scored profiles, with per-criterion coverage across the whole shortlist. - **New**: Clicking a cell filters the stage to the candidates who evidenced that criterion. - **Improved**: Criteria you marked as high importance are ordered first in every view that lists them, not just the review drawer. - **Fixed**: Stages with more than 800 profiles no longer time out when a criterion is edited and the stage is re-run. --- # Arbi v2.17: Reply detection that understands a no > Out of office is not a reply, and neither is a polite decline that you still want counted separately. Sequences now classify what came back instead of only noticing that something did. - Released: 2026-07-30 - Area: Sequencing - Canonical: https://neuroscale.ai/releases/2.17 ## Changes - **New**: Replies are classified as interested, declined, referred, or automatic, and each one stops the sequence differently. - **New**: A decline can write a follow-up date, which puts the candidate back in your queue rather than closing them out. - **Improved**: Out of office replies no longer count against a step's reply rate in analytics. - **Fixed**: Threads with more than twenty messages now render the full history in the inbox instead of truncating at the tenth. --- # Arbi v2.16: Scorecards that fill themselves in > The interviewer's job is to agree or disagree with evidence, not to reconstruct an hour from four lines of handwriting. Recorded interviews now arrive as a scorecard with the passages already attached to each competency. - Released: 2026-07-16 - Area: Interviewing - Canonical: https://neuroscale.ai/releases/2.16 ## Changes - **New**: Competencies are matched to timestamped passages from the transcript, and every rating links back to what was said. - **New**: Interviewers can flag a competency as not covered, which surfaces it as a question for the next round. - **Improved**: Scorecard templates can be shared across a loop so every interviewer rates against the same anchors. --- # Arbi v2.15: Searches that remember what you meant > A good search brief takes real thought to write, and until now it evaporated the moment you closed the tab. Searches are saved objects that keep running. - Released: 2026-07-02 - Area: Sourcing - Canonical: https://neuroscale.ai/releases/2.15 ## Changes - **New**: Any search can be saved, named, and shared with the rest of your team. - **New**: Saved searches re-run on a schedule and show how many candidates are new since you last looked. - **Improved**: Hard filters like work authorisation and location are now separated from the description, so widening the brief does not quietly drop a constraint. - **Fixed**: Duplicate profiles surfaced from more than one source are collapsed into a single result. --- # Arbi v2.14: Every verdict shows its passage > A score nobody can audit is a rumour with a number attached. Each requirement in the review drawer now opens onto the exact sentence in the profile that settled it. - Released: 2026-06-18 - Area: Screening - Canonical: https://neuroscale.ai/releases/2.14 ## Changes - **New**: Clicking a requirement highlights the source passage in place, with the surrounding paragraph left visible for context. - **New**: Not evidenced is now a distinct verdict from a fail, and it renders differently everywhere it appears. - **Improved**: Marking a verdict wrong captures the criterion and the passage as a labelled example, which is what we evaluate model changes against. --- # Arbi v2.13: Analytics per step, not per campaign > A campaign-level reply rate tells you that something is wrong without telling you which message caused it. Performance is now broken out by step. - Released: 2026-06-04 - Area: Sequencing - Canonical: https://neuroscale.ai/releases/2.13 ## Changes - **New**: Open, reply, and decline rates are reported for each step in a sequence. - **New**: Steps can be reordered or removed while a sequence is live without resetting anyone already partway through it. - **Improved**: Sending windows respect the recipient's timezone rather than the sender's. --- # Arbi v2.12: Candidates book their own interviews > The scheduling thread is the slowest part of a fast loop. Candidates now get a portal that shows real availability across every interviewer on the panel. - Released: 2026-05-21 - Area: Interviewing - Canonical: https://neuroscale.ai/releases/2.12 ## Changes - **New**: A candidate portal with live panel availability, rescheduling, and a calendar invite that lands with the right joining details. - **Improved**: Panel conflicts are resolved before times are offered, so a slot cannot be taken twice. - **Fixed**: Invites sent to candidates in a different timezone no longer display the interviewer's local time. --- # Arbi v2.11: SSO, SCIM, and audit exports > The work that makes a tool adoptable by a company rather than a team. Directory sync, enforced sign-on, and a complete record of who decided what. - Released: 2026-05-07 - Area: Platform - Canonical: https://neuroscale.ai/releases/2.11 ## Changes - **New**: SAML single sign-on with enforced login, plus SCIM provisioning against Okta, Entra, and Google Workspace. - **New**: Audit exports covering every verdict, criterion edit, and stage re-run, delivered as CSV or to an S3 bucket you own. - **Improved**: Roles are now scoped per job rather than per workspace, so an agency partner can be given one req and nothing else. - **Fixed**: Deactivating a user no longer detaches the verdicts they recorded from the audit trail. ---