Inference Engineer, vLLM & CUDA
Halide Labs · Remote (EU / US)
$180k–$240k
Ines Duarte
Built exclusively for companies hiring in AI. Discover engineers, researchers, and builders through intelligent matching—not resume databases.
Every candidate is vetted by someone who's actually built and shipped AI systems — not a keyword filter.
Halide Labs · Remote (EU / US)
$180k–$240k
Ines Duarte

Candidate 4 of 7
“Doesn't just build demos — ships systems that hold up in production. Found a reliability bug in our pipeline our own team had missed for weeks.”
Ines Duarte
Member of Technical Staff · Halide Labs
The layer under the automation.
Ines Duarte
Member of Technical Staff, Halide Labs
The problem
Most companies hiring for AI don't have the in-house knowledge to tell who's actually good at it — so they send existing engineers to a course and hope. That's not how this works. New systems need people who've already built and shipped them, and a hiring process wasn't built to find those people. We're not another database. We're your hiring partner: we learn your company and your vision, then find the people who can actually get you there.

We've recently joined hands with their team and we're building a sustainable tech ecosystem together. It's been amazing working with their team — the coordination and execution is simple and efficient, and their talent is brilliant.
How it works
Five steps. The only automated part is the email that tells you which step you're on.
Just the basics — no essay, no résumé formatting. A few details about you and what you're looking for.
Apply to roles on the board, or email hire@0thlayer.com if you're after something specific and want to talk to our team directly.
When we find a role that fits, we reach back out to you before anything else moves forward.
Your application goes to the hiring team. Once they give the green light, we set up your interview.
Once the interview process wraps up, we take care of the documentation and you start your new role.
Ines Duarte
reviewer
Halide Labs · Remote (EU / US)
$180k–$240k
8–64 GPUs
Ines Duarte
Halide Labs
Cadence-9 · London · Hybrid
£145k–£190k
>1k accelerators
Marcus Oyelaran
Cadence-9
Northrail · Remote (US)
$200k–$265k
Single-node
Wei Tan
Northrail
Verge Robotics · Zürich
€155k–€205k
Multi-node
Sofia Marchetti
Verge Robotics
Trellis Health · Remote (US) · Boston
$160k–$205k
Petabyte-scale
Daniel Okonkwo
Trellis Health
Kestrel · Remote (EU)
€120k–€158k
Single-node
Anna Lindqvist
Kestrel
Halide Labs · Remote (EU / US)
Ines Duarte
Halide Labs · Remote (EU / US) · Full-time
You would own the serving path for our 8B and 70B models end to end — the scheduler, the KV-cache, the CUDA kernels underneath it, and the eval harness that tells us when a change made things worse. That path serves about 40 million requests a week across 8 to 64 GPUs, and its p99 is the number the whole company watches on a Monday.
The first thing you would touch is the eviction policy. We already know it is wrong; there is a thread about it on the forum and a graph nobody likes. We would rather hire the person who argues with us in that thread than the person who agrees with the job post.
Ines Duarte
Member of Technical Staff, Halide Labs
We were evicting the least-recently-used block in the pool, which sounds like the neutral choice and is not. The blocks that go longest without a touch belong to the requests that are still decoding — the long ones, the ones with the most prefill already sunk into them. So the policy quietly targets exactly the work you least want to throw away.
Mean latency did not move. p99 tripled inside forty minutes. Every eviction there costs a full prefill replay, and that replay queues behind the arrivals that caused the eviction in the first place. Graph below; the marker is where we noticed.
1# LRU across the whole pool. Looks fair. Is not.2def evict(self, need: int) -> None:3 freed = 04 for block in sorted(self.blocks, key=lambda b: b.last_used):5 if block.pinned:6 continue7 freed += self.release(block) # 16 tokens per block8 if freed >= need:9 return10 raise OutOfCache("pool exhausted")Rollout at −13m. Mean held. The tail did not.
Marcus Oyelaran
Staff Engineer, Training Infrastructure, Cadence-9412 useful · 2h agoLRU on a KV-cache is a scheduling policy wearing a memory policy’s clothes. You are not choosing which bytes are cold, you are choosing which request to punish, and you picked the one with the deepest sunk cost. The tail is the replay, not the miss.
8 marked usefulWei Tan
Research Engineer, Northrail197 useful · 1h agoBefore you rewrite the policy, plot evictions-per-request against sequence length. We had this exact shape and it turned out 4% of requests were absorbing 60% of the evictions. Once that was on screen the fix was an admission check at the front door, not a cleverer rule at the back.
Sofia Marchetti
Principal ML Engineer, Verge Robotics341 useful · 20m agoSeconding the replay theory. Cheap mitigation while you rewrite: pin the blocks of any request that is already past its p50 decode length. Costs a few percent of the pool. Our p99 was back under 210ms within a day and stayed there.
Halide Labs · Remote (EU / US) · closes in 4d
Ines Duarte
Member of Technical Staff · Halide Labs
1d ago
Reviewer is accountable for having read all 213 in full.
I've been working with them for a few years and they've supported my brand a lot — in tech operations, support, and digital activities. They've been our long-term support partner.
The forum
We're not live yet — we're building the forum into a real home for AI practitioners everywhere to think out loud, and for hiring teams to see how they think. Join the waitlist and help shape it from day one.
Ines Duarte
Member of Technical Staff, Halide Labs
We were evicting the least-recently-used block in the pool, which sounds like the neutral choice and is not. The blocks that go longest without a touch belong to the requests that are still decoding — the long ones, the ones with the most prefill already sunk into them. So the policy quietly targets exactly the work you least want to throw away.
Mean latency did not move. p99 tripled inside forty minutes. Every eviction there costs a full prefill replay, and that replay queues behind the arrivals that caused the eviction in the first place. Graph below; the marker is where we noticed.
1# LRU across the whole pool. Looks fair. Is not.2def evict(self, need: int) -> None:3 freed = 04 for block in sorted(self.blocks, key=lambda b: b.last_used):5 if block.pinned:6 continue7 freed += self.release(block) # 16 tokens per block8 if freed >= need:9 return10 raise OutOfCache("pool exhausted")Rollout at −13m. Mean held. The tail did not.
Marcus Oyelaran
Staff Engineer, Training Infrastructure, Cadence-9412 useful · 2h agoLRU on a KV-cache is a scheduling policy wearing a memory policy’s clothes. You are not choosing which bytes are cold, you are choosing which request to punish, and you picked the one with the deepest sunk cost. The tail is the replay, not the miss.
8 marked usefulWei Tan
Research Engineer, Northrail197 useful · 1h agoBefore you rewrite the policy, plot evictions-per-request against sequence length. We had this exact shape and it turned out 4% of requests were absorbing 60% of the evictions. Once that was on screen the fix was an admission check at the front door, not a cleverer rule at the back.
Sofia Marchetti
Principal ML Engineer, Verge Robotics341 useful · 20m agoSeconding the replay theory. Cheap mitigation while you rewrite: pin the blocks of any request that is already past its p50 decode length. Costs a few percent of the pool. Our p99 was back under 210ms within a day and stayed there.
Long posts with code, plots, and numbers — not hot takes. Threaded replies, ordered by time, nothing boosted for engagement. No feed, ever.
Reputation comes from people in your field marking an answer useful — not from posting volume. It'll be on your profile, visible to every hiring team on the platform.
Point at something you've written and say “start here.” It becomes part of how a hiring team gets to know you — not just your résumé.
Wherever you're building AI — inside a lab, a startup, or on your own — this is a place to think out loud with people doing the same work.
We needed engineers who'd actually shipped AI systems before, not just interviewed well. 0th Layer matched us with people who could hit the ground running on our internal AI and automation work — the vetting alone saved us months of bad interviews.
Two sides, one rule
Candidates get matched by people who understand the work. Companies get a hiring partner, not a database. Same team, same process, both directions.
For candidates
Engineers, researchers, infra
For companies
Founders, ML leads, hiring managers

What's on the board
We turn down more listings than we run. If it's “AI-adjacent,” it's somewhere else.
Hiring for AI used to mean wading through hundreds of resumes hoping someone was real. Through 0th Layer we found people who'd genuinely built this stuff before, and they've been building out our internal AI and automation ever since.
We hired the engineers who now run our internal AI tooling and automation through 0th Layer. Every one of them had actually built the kind of systems we needed — that's rare, and it's why we keep coming back to hire more.

Where the line is
Short list. Worth reading before you sign up.
AI ranks and identifies fit across every application, but no model has the final say. A real expert reviews every shortlist before it reaches you.
We don't sell, license, or “enrich” your data into sourcing tools. A company sees your application when you apply to that company. That's the whole distribution list.
Forum posts are public — that's the point of a forum. Your applications, your links, and the fact that you're looking are not. Not to your employer. Not to anyone.
Delete the account and it's gone, threads included if you want them gone. Export first if you'd rather keep it.
Questions about who reads what, what it costs, and what happens to your data. If yours isn't here, email us at hire@0thlayer.com.
Free for candidates, permanently. Companies pay per role. No sourcing spam in either direction.