> This is the markdown version of https://www.classet.ai/blog/ai-recruiting-assistant-risks
> Learn more at https://www.classet.ai



![Risks and evaluation criteria for AI recruiting assistants](/_next/image?url=%2Fimages%2Fblog%2Fai-recruiting-assistant-risks.png&w=3840&q=75)

# AI Recruiting Assistant Guide: Tools, Tips & Risks

The tools work. Most deployments still underperform, and the reasons are predictable. Six failure modes, the compliance line that should decide your shortlist, and a 30-day pilot that produces a real answer.

[![Paul Jones](/_next/image?url=https%3A%2F%2Fassets.basehub.com%2Fe0b5701f%2F6599306507912123f90f150a8bfaaf6c%2Fscreenshot-2026-01-28-at-10.53.16-am.png%3Fwidth%3D100%26height%3D100%26quality%3D100&w=96&q=75)

Paul JonesHead of Growth at Classet

](/blog/authors/paul-jones)

July 9, 2026

AI Recruiting, Guides & Insights

The category works. That part is settled. Jabarian and Henkel's [natural field experiment across roughly 70,000 job candidates](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5395709), run by economists at Chicago Booth and Erasmus Rotterdam, found candidates interviewed by AI were 12% more likely to receive an offer and 18% more likely to still be employed a month later. That's not a vendor case study.

And yet a lot of rollouts land flat. Not because the technology failed, but because the deployment did. The failure modes repeat with enough consistency that you can screen for them before signing, which is what this guide is for. For the ground-level explanation of what these tools do and how they fit an existing process, start with our [AI recruiting assistant guide for TA teams](/blog/ai-recruiting-assistant-ta-leaders). This piece is about what goes wrong.

**TLDR:**

-   Scoring and ranking is the line that determines your compliance obligations, so decide where you stand on it before you shortlist vendors
-   Invitation-based tools show strong results on the candidates who respond, and quietly lose the ones who don't
-   Integrations that read from your ATS but don't write back leave recruiters in two systems and your reporting incomplete
-   Structured summaries need to be checkable against a transcript, or you're making decisions on generated text nobody verified
-   Pilots that measure recruiter satisfaction instead of completion rate and time-to-first-contact produce no decision
-   Run 30 days on one high-volume role family, hold everything else constant, and measure four numbers

## The Six Ways These Rollouts Fail

**1\. Silent rejection.** A tool that auto-rejects candidates below a threshold will eventually reject someone it shouldn't, and nobody will notice because there's no artifact to review. This is both the largest legal exposure and the hardest failure to detect. The fix is architectural: the assistant surfaces structured information and pass or fail results against criteria you defined, and a person makes the rejection call.

**2\. Completion collapse on the wrong format.** Invitation-based tools report completion rates among candidates who started the interview, which is a different denominator than the one you care about. If 40% of invited candidates never open the link, your effective screening coverage is 40% lower than the dashboard suggests. Ask every vendor for completion as a percentage of applications received, not of interviews started.

**3\. The read-only integration.** Plenty of tools connect to your ATS in the sense that they pull candidate records out of it. Fewer write completed screens, transcripts, and status changes back into the candidate record. When results live in a separate dashboard, recruiters work in two systems, your ATS reporting stays incomplete, and the pilot fails on adoption rather than on quality.

**4\. Unverifiable summaries.** An AI-generated summary of a conversation is a useful artifact and a genuinely risky one. If a recruiter can't click through from the summary to the exact transcript passage that supports it, the summary is a claim rather than a record. Require the transcript and, for voice tools, the recording. Then spot-check twenty of them in week one.

**5\. Screening a broken job post faster.** If your listing is vague about pay, shift, or location, your applicants are mismatched before the assistant touches them. Automating the screen makes the mismatch visible sooner, which is worth something, but the money is in fixing the post. Teams who skip this conclude the tool didn't work.

**6\. The pilot that never ends.** Rollouts stall when the success criteria were never defined numerically. "Do recruiters like it" produces a six-month evaluation and no decision. Pick the numbers first.

## Tools by What They Actually Automate

Vendors in this category are often compared as if they're interchangeable. They automate different stages, and that determines fit more than any feature list.

Category

What it automates

Format

Watch for

Scheduling and orchestration

Interview coordination, reminders, onboarding steps

Chat and SMS

Doesn't close the apply-to-contact gap on its own

Invited AI interviews

First-round screening on the candidate's schedule

Video, web audio, or phone link

Completion rate against applications received

Outbound AI phone screening

First-round screening initiated by the system

Outbound voice call

Whether the call handles candidate questions, not just asks them

Resume and application parsing

Ranking and filtering inbound volume

Batch scoring

This is the category most likely to trigger AEDT rules

Sourcing and outreach

Finding and messaging passive candidates

Email and InMail

Deliverability and message quality at volume

Most teams need one or two of these, not all five. Mapping your actual bottleneck to a row here eliminates most of the market before you sit through a single demo.

## The Compliance Line That Decides Your Shortlist

There's one question that splits the vendor field cleanly: does the tool score, rank, or filter candidates in a way that substantially assists a hiring decision?

If yes, you're likely in Automated Employment Decision Tool territory under [NYC Local Law 144](https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page) and similar frameworks emerging elsewhere. That means an annual independent bias audit with published results, candidate notice requirements, and documented human oversight. [Deloitte's summary](https://www.deloitte.com/us/en/services/audit-assurance/articles/nyc-local-law-144-algorithmic-bias.html) is a reasonable orientation before you talk to legal.

None of that is disqualifying. Plenty of well-run programs use scoring tools and meet the requirements properly, and we've written about [what a defensible bias audit looks like](/blog/ai-recruiting-bias-audits-compliant-platform). But it changes your procurement work, your timeline, and your ongoing obligations, so you want to know which side of the line a vendor sits on in the first conversation rather than the fourth.

Classet made a deliberate choice here. Joy conducts the interview, applies the knockout criteria you define, and produces a transcript, recording, and structured summary. It does not score or rank. A recruiter reads and decides. Classet also publishes continuous third-party bias audit results through [Warden AI](https://www.warden-ai.com/), and holds SOC 2 Type II certification from Prescient Assurance. If your legal team's first question is "who is accountable for this rejection," that architecture gives you a short answer.

## A 30-Day Pilot That Produces a Real Answer

Most pilots fail as experiments, not as products. Here's a design that doesn't.

Pick one role family with steady volume. Warehouse associate, CNA, service technician, driver. High enough volume that 30 days gives you a real sample, narrow enough that the screening criteria are stable.

Hold everything else constant. Same job posts, same ad spend, same recruiters, same offer process. If you change the job description mid-pilot, you've lost your comparison.

Measure four numbers, before and after:

1.  **Median minutes from application submitted to first candidate contact.** Median, not average, and including nights and weekends.
2.  **Screening completion as a percentage of applications received.** Not of interviews started.
3.  **Recruiter hours per week spent on first-round screening.** Measured for two weeks beforehand, not estimated.
4.  **Apply-to-interview conversion rate.** The share of applicants who reach a human interview.

Then do one thing that isn't a number: sit a recruiter down with twenty completed screens and the matching transcripts, and have them mark where the summary was right, thin, or wrong. Twenty is enough to see a pattern, and it's the check that catches the failure modes a dashboard won't.

Thirty days on one role family, four numbers, one manual review. That produces a decision.

## FAQ

Is using an AI recruiting assistant legal?

Yes, with conditions that depend on where you hire. NYC Local Law 144 requires an annual independent bias audit, published results, and candidate notice for tools that qualify as Automated Employment Decision Tools, meaning they substantially assist hiring decisions through scoring or ranking. Illinois, Maryland, and Colorado have their own requirements. Tools that conduct structured interviews and surface information without scoring generally fall outside AEDT classification, though you still owe candidates notice and should document human oversight.

What's the biggest risk of AI screening?

Auto-rejection without a reviewable artifact. If candidates are being filtered out below a score threshold and no person reviewed the decision, you have both a compliance exposure and a quality problem you can't detect, because rejected candidates don't come back to tell you the tool was wrong. Require that every rejection has a named human decision-maker and a transcript a person could review if challenged.

How do I know if the AI summary is accurate?

Check it against the transcript. Any tool that gives you a summary without the underlying transcript and, for voice tools, the recording, is asking you to trust generated text. In your first pilot week, pull twenty completed screens and have a recruiter compare each summary to the source. If more than a couple are thin or wrong, that's a product problem worth raising before you scale.

Will candidates refuse to talk to an AI?

Some will, and you should offer a path to a human. In practice, refusal rates are lower than most teams expect for frontline roles, and completion often improves over the format it replaces, which is part of [why we don't do video AI interviews](/blog/why-no-video-ai-interviews). When one home services employer moved from a written assessment to a voice screen, completion went from 50% to over 70%, because candidates could talk instead of type. Reporting like NPR's coverage of [AI-conducted job interviews](https://www.npr.org/2025/11/07/nx-s1-5600127/recruiting-companies-are-starting-to-hold-job-interviews-using-ai) shows the reaction is mixed and worth measuring in your own population rather than assuming.

How long should an AI recruiting assistant pilot run?

Thirty days on a single high-volume role family, with everything else held constant. That's long enough to get a real sample on completion rate and time-to-first-contact, and short enough to force a decision. Turnover and quality-of-hire effects take at least two quarters to read, so don't build the pilot around them. Setup shouldn't eat the window either: self-serve tools launch same-day and full ATS integrations typically take two to three weeks.

## Next Step

Run the four numbers on your current process first. Most teams discover their median time-to-first-contact is far worse than the average they've been quoting, and that alone reframes the buying decision.

When you're ready to compare against a live baseline, [book a demo](/demo) and we'll screen candidates on one of your open roles. Our [vendor questions checklist](/blog/ai-recruiting-vendor-questions) covers what to ask before you sign anything.

[![Paul Jones](/_next/image?url=https%3A%2F%2Fassets.basehub.com%2Fe0b5701f%2F6599306507912123f90f150a8bfaaf6c%2Fscreenshot-2026-01-28-at-10.53.16-am.png%3Fwidth%3D100%26height%3D100%26quality%3D100&w=128&q=75)

Paul Jones

Head of Growth at Classet

Paul comes from an operator background running an Alpine-owned company, and brings firsthand experience with the hiring challenges Classet was built to solve. He's driven by a belief that the right technology can make meaningful work more accessible.

](/blog/authors/paul-jones)

Follow Classet in Google

Add classet.ai as a preferred source and our hiring research is more likely to show up in your Top Stories and AI Overviews.

[Add as preferred source →](https://www.google.com/preferences/source?q=classet.ai)

## Explore More

### Use Cases

-   [RPO / BPO Recruiting](/use-cases/call-centers-bpo)
-   [Healthcare Recruiting](/use-cases/healthcare)
-   [Hospitality Recruiting](/use-cases/hospitality)

### Integrations

-   [Greenhouse](/integrations/greenhouse)
-   [Bullhorn](/integrations/bullhorn)
-   [Lever](/integrations/lever)