Skip to content

The AI Talent Layer.

Built exclusively for companies hiring in AI. Discover engineers, researchers, and builders through intelligent matching—not resume databases.

Every candidate is vetted by someone who's actually built and shipped AI systems — not a keyword filter.

3rdYour company
2ndYour product
1stYour AI systems
0thThe talent that builds all of it

What the product produces

Inference Engineer, vLLM & CUDA

Halide Labs · Remote (EU / US)

$180k–$240k

PyTorchTritonCUDA
ID

Reviewed by

Ines Duarte

signed
Someone reading on a laptop alone at a desk late at night, printed pages beside them.
shortlist signed

Candidate 4 of 7

Doesn't just build demos — ships systems that hold up in production. Found a reliability bug in our pipeline our own team had missed for weeks.

Ines Duarte

Member of Technical Staff · Halide Labs

The layer under the automation.

reliabilityperformance20m ago

How fast we catch problems before they cost you

ID

Ines Duarte

Member of Technical Staff, Halide Labs

0120240360480−23m−11mnow
Response time · ms
41 replies12 useful

The problem

You can't retrain your way to an AI team.

Most companies hiring for AI don't have the in-house knowledge to tell who's actually good at it — so they send existing engineers to a course and hope. That's not how this works. New systems need people who've already built and shipped them, and a hiring process wasn't built to find those people. We're not another database. We're your hiring partner: we learn your company and your vision, then find the people who can actually get you there.

Someone part-way through a job application at a cluttered desk in the late afternoon, looking tired rather than distressed.
Good candidate, auto-rejectedKeyword match: 61%Filtered by job title, not skill412 resumes. 3 worth a call.Sourced from a stale databaseStill open after 6 months
We've recently joined hands with their team and we're building a sustainable tech ecosystem together. It's been amazing working with their team — the coordination and execution is simple and efficient, and their talent is brilliant.
Imran

How it works

One human step, put back where it belongs.

Five steps. The only automated part is the email that tells you which step you're on.

  1. 0th

    Create your profile

    Just the basics — no essay, no résumé formatting. A few details about you and what you're looking for.

  2. 1st

    Apply, or reach out directly

    Apply to roles on the board, or email hire@0thlayer.com if you're after something specific and want to talk to our team directly.

  3. 2nd

    We match you to the right role

    When we find a role that fits, we reach back out to you before anything else moves forward.

  4. 3rd

    We submit you and set up the interview

    Your application goes to the hiring team. Once they give the green light, we set up your interview.

  5. 4th

    You're in — we handle the paperwork

    Once the interview process wraps up, we take care of the documentation and you start your new role.

0thlayer.com/jobs
0Halide Labs
BoardSaved12Applications3ForumActivity
ID

Ines Duarte

reviewer

Job board214 open roles · updated 12m ago
inference, CUDA, remote⌘K
stackPyTorchphasescaleremoteEU / US6 of 214
Role CompensationReviewer

Inference Engineer, vLLM & CUDA

signed

Halide Labs · Remote (EU / US)

PyTorchTritonCUDAvLLM

$180k–$240k

8–64 GPUs

ID

Ines Duarte

Halide Labs

Training Infrastructure Engineer

Cadence-9 · London · Hybrid

JAXRay

£145k–£190k

>1k accelerators

MO

Marcus Oyelaran

Cadence-9

Evals Lead, Post-training

signed

Northrail · Remote (US)

PythonPyTorchWeights & Biases

$200k–$265k

Single-node

WT

Wei Tan

Northrail

RL Engineer, Manipulation

Verge Robotics · Zürich

PyTorchMuJoCo

€155k–€205k

Multi-node

SM

Sofia Marchetti

Verge Robotics

ML Data Engineer

Trellis Health · Remote (US) · Boston

SparkdbtAirflowDuckDB

$160k–$205k

Petabyte-scale

DO

Daniel Okonkwo

Trellis Health

Applied ML Engineer, Ranking

Kestrel · Remote (EU)

PyTorchFeast

€120k–€158k

Single-node

AL

Anna Lindqvist

Kestrel

Role detailesc
shortlist signed

Inference Engineer, vLLM & CUDA

Halide Labs · Remote (EU / US)

Comp
$180k–$240k
Scale
8–64 GPUs
Applicants
213
Read
7 of 213
ID

Reviewed by

Ines Duarte

I've been working with them for a few years and they've supported my brand a lot — in tech operations, support, and digital activities. They've been our long-term support partner.
Toshi, Keeper Labo Singapore

The forum

Practitioners worldwide, building this together.

We're not live yet — we're building the forum into a real home for AI practitioners everywhere to think out loud, and for hiring teams to see how they think. Join the waitlist and help shape it from day one.

0thlayer.com/forum/kv-cache-eviction-p99
reliabilityperformance41 replies · last 20m ago

How fast we catch problems before they cost you

ID

Ines Duarte

Member of Technical Staff, Halide Labs

12 marked useful

We were evicting the least-recently-used block in the pool, which sounds like the neutral choice and is not. The blocks that go longest without a touch belong to the requests that are still decoding — the long ones, the ones with the most prefill already sunk into them. So the policy quietly targets exactly the work you least want to throw away.

Mean latency did not move. p99 tripled inside forty minutes. Every eviction there costs a full prefill replay, and that replay queues behind the arrivals that caused the eviction in the first place. Graph below; the marker is where we noticed.

cache/policy.py
1# LRU across the whole pool. Looks fair. Is not.2def evict(self, need: int) -> None:3    freed = 04    for block in sorted(self.blocks, key=lambda b: b.last_used):5        if block.pinned:6            continue7        freed += self.release(block)  # 16 tokens per block8        if freed >= need:9            return10    raise OutOfCache("pool exhausted")
0120240360480−23m−17m−11m−5mnowp99 441 ms
p99 request latency · 24 min window · ms

Rollout at −13m. Mean held. The tail did not.

3 of 41 repliessorted by useful marks
MO

Marcus Oyelaran

Staff Engineer, Training Infrastructure, Cadence-9412 useful · 2h ago

LRU on a KV-cache is a scheduling policy wearing a memory policy’s clothes. You are not choosing which bytes are cold, you are choosing which request to punish, and you picked the one with the deepest sunk cost. The tail is the replay, not the miss.

8 marked useful
WT

Wei Tan

Research Engineer, Northrail197 useful · 1h ago

Before you rewrite the policy, plot evictions-per-request against sequence length. We had this exact shape and it turned out 4% of requests were absorbing 60% of the evictions. Once that was on screen the fix was an admission check at the front door, not a cleverer rule at the back.

SM

Sofia Marchetti

Principal ML Engineer, Verge Robotics341 useful · 20m ago

Seconding the replay theory. Cheap mitigation while you rewrite: pin the blocks of any request that is already past its p50 decode length. Costs a few percent of the pool. Our p99 was back under 210ms within a day and stayed there.

  • Threads, not takes

    Long posts with code, plots, and numbers — not hot takes. Threaded replies, ordered by time, nothing boosted for engagement. No feed, ever.

  • Reputation you earned

    Reputation comes from people in your field marking an answer useful — not from posting volume. It'll be on your profile, visible to every hiring team on the platform.

  • Attach a thread to an application

    Point at something you've written and say “start here.” It becomes part of how a hiring team gets to know you — not just your résumé.

  • Open to practitioners everywhere

    Wherever you're building AI — inside a lab, a startup, or on your own — this is a place to think out loud with people doing the same work.

We needed engineers who'd actually shipped AI systems before, not just interviewed well. 0th Layer matched us with people who could hit the ground running on our internal AI and automation work — the vetting alone saved us months of bad interviews.
Alvin, Advanteq Engineering

Two sides, one rule

Both sides work with the same team.

Candidates get matched by people who understand the work. Companies get a hiring partner, not a database. Same team, same process, both directions.

Get matched, not filtered.

  • A simple profile — not a résumé rewrite
  • We match you to roles that actually fit
  • Real interviews with real people
  • Straight to the hiring team once you're a fit
  • No recruiter spam. We don't have your inbox to sell.

For candidates

Engineers, researchers, infra

Get a hiring partner, not a database.

  • A call to actually understand the role
  • AI ranks every application, a real expert reviews it
  • Candidates matched to your vision, not just keywords
  • We run interviews built around what we learned
  • We stay involved after you sign

For companies

Founders, ML leads, hiring managers

Hiring for AI used to mean wading through hundreds of resumes hoping someone was real. Through 0th Layer we found people who'd genuinely built this stuff before, and they've been building out our internal AI and automation ever since.
Richard, Artist Farm
We hired the engineers who now run our internal AI tooling and automation through 0th Layer. Every one of them had actually built the kind of systems we needed — that's rare, and it's why we keep coming back to hire more.
Suraj Naik, Techflex
A person reading a printed page at a plain desk in an ordinary workroom, lit by a window.

Where the line is

What we don't do.

Short list. Worth reading before you sign up.

  • No AI-only decisions.

    AI ranks and identifies fit across every application, but no model has the final say. A real expert reviews every shortlist before it reaches you.

  • No selling your profile.

    We don't sell, license, or “enrich” your data into sourcing tools. A company sees your application when you apply to that company. That's the whole distribution list.

  • No profile you didn't choose to make public.

    Forum posts are public — that's the point of a forum. Your applications, your links, and the fact that you're looking are not. Not to your employer. Not to anyone.

  • No lock-in.

    Delete the account and it's gone, threads included if you want them gone. Export first if you'd rather keep it.

FAQs

Questions about who reads what, what it costs, and what happens to your data. If yours isn't here, email us at hire@0thlayer.com.

One application. One person. One answer.

Free for candidates, permanently. Companies pay per role. No sourcing spam in either direction.