TLDR
Surveyed ~40 openings across article on-ramps + live boards (seen 2026-09-15). Ranking weights: remote-US friendliness × applied eng / red-team / eval / control / shipping agents × safety-impact credibility × realistic odds for a non-PhD career switcher × comp + stability. Deprioritized: UK-only, pure research-scientist/PhD-pubs tracks, heavy onsite SF with no remote, unpaid unless strong on-ramp.
Top 5 apply-now: (1) METR Task Development Engineer · (2) FAR.AI Research Engineer / Senior RE · (3) FAR.AI Jailbreaking Lead, Red Team · (4) NVIDIA Evaluation & ML Systems Engineer (AI Safety & Security) — confirm TX eligibility · (5) Epoch AI Researcher (esp. benchmarking).
Public profile used only: MS CS; years shipping applied AI; daily multi-agent eng across frontier + open models; Anthropic model-safety bounty enrolled with a harness built; builder/red-team/eval energy; family of seven in Tyler TX area; remote-US preferred.
How this list was ranked
- Remote-US / timezone: America/Chicago preferred; rare travel OK; “SF or nothing” only if exceptional fit+pay.
- Fit: harnesses, agent evals, red-team/disclosure, control/monitoring, shipping — not pure mech-interp theory.
- Impact credibility: independent evaluators & frontier safety eng ahead of vague “AI ethics” titles.
- Odds: eng-adjacent > PhD-pubs research scientist roles.
- Comp/stability: posted bands preferred; ESTIMATEs cited to Levels.fyi / 80k Hours / Glassdoor-class sources carefully.
Honesty: several lab career pages are JS-heavy; some titles come via Greenhouse/Lever/Ashby/80,000 Hours mirrors checked today. Closed/expired flagged when known. Do not invent precision on pay.
#1 · Task Development Engineer · METR
Why for Eric: Closest FTE/contract match to daily multi-agent harness + eval work. METR’s Time Horizons / agent capability evals are high-credibility independent safety measurement — and they hire remote with Pacific overlap (Chicago mornings work).
Requirements (public): Several years complex software eng; experience building hard (ideally agent-based) AI evals (RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA); Inspect preferred; Hawk / Time Horizons nice-to-have; high attention to detail.
Logistics: Remote worldwide; ~1–4h overlap with Pacific workday; 20–40h/wk flexible on contractor framing. America/Chicago is fine for overlap.
Travel: Not required for remote contractor framing; FTE path mentions Berkeley office culture / work trials — confirm.
Pay: Posted (careers / 80k Hours): $150–$300/hour contractor. Also seen on Lever (2026): FTE band $260,937–$385,490/year with Bay benefits/relocation language. Reconcile at apply time — page text recently shifted toward FTE-by-default while careers card still says remote contractor.
Day-to-day: Design novel hard tasks as model horizons grow; QA solvability; baseline/score; improve task-dev infrastructure.
Link: metr.org/careers · Lever posting
Caveats: Competitive; eval portfolio helps more than titles. Contract vs FTE ambiguity. Not a “lab insider” seat — independent evaluator impact.
#2 · Research Engineer / Senior Research Engineer · FAR.AI
Why for Eric: Explicitly ships pre/post-release adversarial evals of frontier models, novel attacks, robustness — builder/shipper safety without PhD-first framing. Remote + Berkeley optional.
Requirements (public, Research Engineer page historically): Implement ML algorithms, run experiments, analyze results; evals/red-teaming; open-source ML stack familiarity (PyTorch/HF). Senior bar = more ownership/scope.
Logistics: Remote and in-person Berkeley possible; hire remotely in most countries (80k Hours: Remote, global).
Travel: Work-related travel/equipment covered per older posting language; not “SF or nothing.”
Pay: Posted via 80,000 Hours (seen 2026-09-15): Senior Research Engineer $150–$250k. Research Engineer band not always listed on the card — treat as unknown / ask (older FAR page cited wide location-dependent ranges; do not invent).
Day-to-day: Frontier model adversarial evals; attack development; experiments on deception/robustness; code + paper co-authorship as MTS-style eng.
Links: far.ai/careers · RE Ashby 52e76732… · Senior RE 4f6fece8…
Caveats: Nonprofit/research org scale vs frontier-lab cash; Ashby pages are JS shells — read full JD in-browser. Competition still real among alignment-adjacent engineers.
#3 · Jailbreaking Lead, Red Team · FAR.AI
Why for Eric: Direct line from enrolled Anthropic model-safety bounty + harness building into a paid red-team leadership seat at an org whose red team publishes jailbreak/finetune attack research affecting Claude/ChatGPT/Gemini safeguards.
Requirements: Mid (5–9y) experience per 80k Hours card; deep adversarial / jailbreak / eval execution (confirm full JD on Ashby).
Logistics: Remote, global.
Travel: Unknown beyond normal research org trips — assume low vs lab hybrid.
Pay: Posted (80k Hours): $170–$250k.
Day-to-day: Lead jailbreak/red-team technical strategy; evaluate frontier systems; turn findings into research + safeguard improvements.
Link: Ashby Jailbreaking Lead · board mirror 80k Hours · FAR AI
Caveats: “Lead” title may expect prior published attacks or team lead proof — portfolio of authorized red-team results matters. Not the same as product pentest-only background.
#4 · Evaluation & ML Systems Engineer · NVIDIA (AI Safety & Security)
Why for Eric: Evidence-first eval infrastructure for AI-powered vuln find/validate/patch — maps to shipping harnesses, metrics skepticism, agent measurement. Big-company stability + safety-adjacent cyber.
Requirements (public): Bachelor’s or equivalent + 5+ years ML eng/evaluation; designing benchmarks/metrics; solid Python for shared infra; experiment tracking/data pipelines. Preferred: security evaluation, agent/LLM behavior measurement, public benchmarks.
Logistics: US remote in listed locations (mirrors commonly: Santa Clara + remote CA/NC/NY/TN/FL). Texas not consistently listed on this exact JD — verify on jobs.nvidia.com before prioritizing.
Travel: Unknown; typical corp remote.
Pay: Posted base: $152,000–$241,500. ESTIMATE total comp: NVIDIA ML Eng Levels.fyi US medians often ~$200k–$330k+ TC depending on level (stock-heavy) — not this JD’s offer.
Day-to-day: Build benchmarking/reproducibility systems; define metrics/protocols; map every result to code+runs; keep conclusions reviewable.
Link: NVIDIA job 893396714449
Caveats: State eligibility may block Tyler TX; applications may already be past “accept until July 30, 2026” language on some mirrors — confirm still open. Impact is AI-for-security tooling more than frontier alignment theory.
#5 · Researcher / Senior Researcher · Epoch AI
Why for Eric: Fully remote with PT–CET hiring; measurement/critique of AI progress fits “make numbers mean something.” Lower day-to-day red-team than #1–#3, but family-logistics excellent and credibility high.
Requirements: Open to varied backgrounds; researcher roles across multiple teams (incl. Benchmarking Reviews producing critiques of AI benchmarks). Prefer overlap with PT and UTC; can travel to ~3 retreats/year.
Logistics: Fully remote; many countries PT→CET. America/Chicago OK.
Travel: Prefer candidates who can attend ~3 staff retreats/year.
Pay: Posted examples: Researcher/Senior across teams often cited in wide bands (e.g. Benchmarking Reviews mirrors $100k–$200k); broader researcher postings sometimes higher — confirm per team. Do not assume top of range.
Day-to-day: Research trends/capabilities; write reviews/critiques; publish for policymakers and industry.
Links: Epoch Lever researcher · epoch.ai careers
Caveats: More research writing than agent harness shipping; retreat travel; not adversarial red-team primary.
#6 · Research Engineer, Model Evaluations · Anthropic
Why for Eric: Design/run capability + safety evals; harden distributed eval platforms against training checkpoints; dashboards under time pressure — almost a job description of “ship eval harnesses.” Lab-scale impact.
Requirements: Strong Python; distributed systems/data pipelines; clear communication; on-call comfort during training runs. Preferred: LLM scaffolding, dashboards, eval metrics, stats, ML infra.
Logistics: SF or NYC; hybrid ≥25% in office; labeled Remote-Friendly (Travel-Required).
Travel: Material — plan regular SF/NY presence from Texas.
Pay: Posted: $500,000–$850,000 annual salary (Greenhouse).
Day-to-day: New evals (reasoning, agents, safety); distributed eval execution; mid-run debugging; researcher partnership.
Link: Greenhouse 5198255008
Caveats: Family travel cost is the real filter. Extremely competitive. Visa sponsorship possible but not guaranteed for every candidate.
#7 · Red Team Engineer, Safeguards · Anthropic
Why for Eric: Jailbreaks, agentic prompt-injection, automated testing frameworks — closest lab red-team title to bounty harness work.
Requirements: Pentest/red-team/appsec experience; model jailbreaking; agentic workflow testing; AI/ML security preferred; API/authz abuse, T&S backgrounds helpful.
Logistics: SF; Remote-Friendly (Travel Required); ≥25% office policy.
Travel: Yes — hybrid expectation.
Pay: Not clearly listed on the public card fetched today → unknown (ask). Sibling Anthropic RE roles post very high bands; do not invent a number.
Day-to-day: Creative multi-step attacks on product surfaces; systematic methodologies; automated LLM testing frameworks.
Link: Greenhouse 5320469008
Caveats: Hybrid travel; may prefer classical security pedigrees alongside LLM jailbreak skill. Competition high.
#8 · Engineering Manager, Red Team · FAR.AI
Why for Eric: If multi-agent fleet coordination / Board-style review generalizes to leading a red-team eng org — remote path into FAR red-team. Skip if you want IC-only.
Requirements: Mid experience managing eng delivery (confirm Ashby JD).
Logistics / travel: Remote; travel unknown.
Pay: Posted (80k Hours): $170–$250k.
Day-to-day: People/project systems enabling red-team research velocity.
Link: Ashby EM Red Team
Caveats: Manager ≠ IC red-teamer; hiring committees weight management receipts.
#9 · Research Engineer, Frontier AI Risk Management · SaferAI
Why for Eric: Translate risk commitments into concrete evals/mitigations/red-teams of mitigations — eng-adjacent assurance work.
Requirements: Technical RE skills for evals/mitigations; async remote collaboration; strong writing.
Logistics: Team mostly Paris/London; remote for particularly strong candidates.
Travel / pay: Unknown. Pay: unknown (not posted on page fetched).
Link: safer-ai.org jobs (FAIRM Research Engineer) — verify live; some mirrors 404’d on fetch.
Caveats: EU timezone bias; page may close when filled; weaker US remote guarantee than METR/FAR/Epoch.
#10 · MTS, Safety for Agents · Cohere
Why for Eric: Safety for agents that take actions — data gen, post-training, evals. Multi-agent shipping experience is relevant.
Requirements: Strong SE + stats/experimental design; data collection with annotators; Python/ML frameworks.
Logistics: Remote-friendly across NA/Europe/etc.; ~50% working-day overlap with UK/EU (US East fine; Chicago is early mornings).
Pay: Third-party mirrors have cited CAD bands — treat as ESTIMATE / unverified here; confirm on Cohere careers.
Day-to-day: Safety post-training + eval methods for tool-using models; cross-functional with product/policy.
Caveats: Timezone tax from Texas; confirm role still open.
Honorable mentions (logistics hard)
- Apollo Research — AI Red Team Engineer: Superb scheming/control/monitor red-team fit (Watcher). In-person London or SF with flexible WFH. Posted SF $182k–$238k / London £122–160k. Lever
- Redwood Research — Member of Technical Staff: AI control, alignment-faking, safety cases. Berkeley onsite. Posted $350k–$850k. careers
- Anthropic — Cyber Evaluations Engineer: Cyber capability + safeguard robustness evals; SF | DC; hybrid policy. Greenhouse
- OpenAI Preparedness (Automated Red Teaming, Threat Modeler, etc.): high fit, San Francisco. ART example posted ~$295k–$445k.
- xAI / SpaceX AI safety eng: Palo Alto / office-heavy (e.g. Imagine Safety; Safety Security).
- METR MTS (Eval Execution, Embedded Assessments, Security Eng, Cyberforensics): Berkeley onsite/hybrid; posted roughly $328k–$687k depending on seat.
- NIST CAISI: US citizenship; in-person DC & SF; Research Engineer/Scientist and MTS hiring — not remote-first. nist.gov/caisi
- UK AISI: Strong red-team/control openings (£65–145k bands) — London / UK, visa. Closes noted on some RE roles ~30 Sep 2026.
On-ramps / not FTE jobs
- BlueDot Impact — Technical AI Safety (free course): bluedot.org
- MATS research fellowships: matsprogram.org
- ARENA technical bootcamps: arena.education
- Anthropic Fellows (AI Safety / AI Security): 4 months FT; remote-friendly US/UK/Canada with work auth; stipend $3,850/week USD (+ ~$15k/mo compute). On-ramp with historically strong conversion to Anthropic safety FTE. Fellows Greenhouse
- Anthropic model safety bug bounty + red.anthropic.com — already enrolled; finish a clean claim.
- xAI/SpaceX HackerOne (authorized only): hackerone.com/x
- 80,000 Hours map + board: 80000hours.org/ai · jobs.80000hours.org
- METR General Expression of Interest (flexible): metr.org/careers
- Hawk open eval stack: hawk.metr.org — practice Inspect-like task work publicly.
Suggested apply order (family-aware)
- METR Task Development Engineer + FAR RE / Jailbreaking Lead (true remote).
- Epoch researcher + NVIDIA eval eng if TX (or relocate-state) eligibility clears.
- Anthropic Fellows (time-boxed on-ramp) or Anthropic Model Evals / Safeguards Red Team if hybrid math works.
- Keep Apollo/Redwood/OpenAI/xAI/CAISI as stretch only if relocation or rare-travel exceptions appear.
Parallel portfolio (1–2 weeks): one authorized bounty submission; one public reproduction of a METR/Inspect-style eval with a clean negative-or-partial result write-up; link Hawk familiarity.
Sources & honesty
Checked 2026-09-15 (UTC) via WebFetch/WebSearch/curl: metr.org/careers + Lever; far.ai / Ashby / jobs.80000hours.org; Anthropic Greenhouse; NVIDIA job page mirrors; Epoch Lever; Apollo Lever + careers FAQ; Redwood careers; NIST CAISI; UK AISI EOIs; OpenAI/xAI search results; article on-ramps in /workspace/eric-jobs/labs.txt & long.txt.
Machine-readable summary: /workspace/eric-jobs/ranked-jobs.json.
No secrets. No fabricated private résumé bullets. If a posting closed between research and your click, trust the live page.