← Digests

Titus digest

Remote-friendly AI-safety jobs pack

2026-09-15 · ranked for Eric Buess (America/Chicago, applied multi-agent eng, red-team/evals/control preference) · from Grok Bot

Voice picker uses whatever Safari exposes. Jamie Premium in iOS Spoken Content often does not appear here — that’s an Apple limit. If missing, pick another voice or use pre-recorded audio.

1 / 16
Ready.

TLDR

Surveyed ~40 openings across article on-ramps + live boards (seen 2026-09-15). Ranking weights: remote-US friendliness × applied eng / red-team / eval / control / shipping agents × safety-impact credibility × realistic odds for a non-PhD career switcher × comp + stability. Deprioritized: UK-only, pure research-scientist/PhD-pubs tracks, heavy onsite SF with no remote, unpaid unless strong on-ramp.

Top 5 apply-now: (1) METR Task Development Engineer · (2) FAR.AI Research Engineer / Senior RE · (3) FAR.AI Jailbreaking Lead, Red Team · (4) NVIDIA Evaluation & ML Systems Engineer (AI Safety & Security) — confirm TX eligibility · (5) Epoch AI Researcher (esp. benchmarking).

Public profile used only: MS CS; years shipping applied AI; daily multi-agent eng across frontier + open models; Anthropic model-safety bounty enrolled with a harness built; builder/red-team/eval energy; family of seven in Tyler TX area; remote-US preferred.

How this list was ranked

  • Remote-US / timezone: America/Chicago preferred; rare travel OK; “SF or nothing” only if exceptional fit+pay.
  • Fit: harnesses, agent evals, red-team/disclosure, control/monitoring, shipping — not pure mech-interp theory.
  • Impact credibility: independent evaluators & frontier safety eng ahead of vague “AI ethics” titles.
  • Odds: eng-adjacent > PhD-pubs research scientist roles.
  • Comp/stability: posted bands preferred; ESTIMATEs cited to Levels.fyi / 80k Hours / Glassdoor-class sources carefully.

Honesty: several lab career pages are JS-heavy; some titles come via Greenhouse/Lever/Ashby/80,000 Hours mirrors checked today. Closed/expired flagged when known. Do not invent precision on pay.

#1 · Task Development Engineer · METR

remote worldwide evals / agents top fit

Why for Eric: Closest FTE/contract match to daily multi-agent harness + eval work. METR’s Time Horizons / agent capability evals are high-credibility independent safety measurement — and they hire remote with Pacific overlap (Chicago mornings work).

Requirements (public): Several years complex software eng; experience building hard (ideally agent-based) AI evals (RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA); Inspect preferred; Hawk / Time Horizons nice-to-have; high attention to detail.

Logistics: Remote worldwide; ~1–4h overlap with Pacific workday; 20–40h/wk flexible on contractor framing. America/Chicago is fine for overlap.

Travel: Not required for remote contractor framing; FTE path mentions Berkeley office culture / work trials — confirm.

Pay: Posted (careers / 80k Hours): $150–$300/hour contractor. Also seen on Lever (2026): FTE band $260,937–$385,490/year with Bay benefits/relocation language. Reconcile at apply time — page text recently shifted toward FTE-by-default while careers card still says remote contractor.

Day-to-day: Design novel hard tasks as model horizons grow; QA solvability; baseline/score; improve task-dev infrastructure.

Link: metr.org/careers · Lever posting

Caveats: Competitive; eval portfolio helps more than titles. Contract vs FTE ambiguity. Not a “lab insider” seat — independent evaluator impact.

#2 · Research Engineer / Senior Research Engineer · FAR.AI

remote global evals + red-team

Why for Eric: Explicitly ships pre/post-release adversarial evals of frontier models, novel attacks, robustness — builder/shipper safety without PhD-first framing. Remote + Berkeley optional.

Requirements (public, Research Engineer page historically): Implement ML algorithms, run experiments, analyze results; evals/red-teaming; open-source ML stack familiarity (PyTorch/HF). Senior bar = more ownership/scope.

Logistics: Remote and in-person Berkeley possible; hire remotely in most countries (80k Hours: Remote, global).

Travel: Work-related travel/equipment covered per older posting language; not “SF or nothing.”

Pay: Posted via 80,000 Hours (seen 2026-09-15): Senior Research Engineer $150–$250k. Research Engineer band not always listed on the card — treat as unknown / ask (older FAR page cited wide location-dependent ranges; do not invent).

Day-to-day: Frontier model adversarial evals; attack development; experiments on deception/robustness; code + paper co-authorship as MTS-style eng.

Links: far.ai/careers · RE Ashby 52e76732… · Senior RE 4f6fece8…

Caveats: Nonprofit/research org scale vs frontier-lab cash; Ashby pages are JS shells — read full JD in-browser. Competition still real among alignment-adjacent engineers.

#3 · Jailbreaking Lead, Red Team · FAR.AI

remote global red-team lead

Why for Eric: Direct line from enrolled Anthropic model-safety bounty + harness building into a paid red-team leadership seat at an org whose red team publishes jailbreak/finetune attack research affecting Claude/ChatGPT/Gemini safeguards.

Requirements: Mid (5–9y) experience per 80k Hours card; deep adversarial / jailbreak / eval execution (confirm full JD on Ashby).

Logistics: Remote, global.

Travel: Unknown beyond normal research org trips — assume low vs lab hybrid.

Pay: Posted (80k Hours): $170–$250k.

Day-to-day: Lead jailbreak/red-team technical strategy; evaluate frontier systems; turn findings into research + safeguard improvements.

Link: Ashby Jailbreaking Lead · board mirror 80k Hours · FAR AI

Caveats: “Lead” title may expect prior published attacks or team lead proof — portfolio of authorized red-team results matters. Not the same as product pentest-only background.

#4 · Evaluation & ML Systems Engineer · NVIDIA (AI Safety & Security)

US remote (state list) evals / security tooling

Why for Eric: Evidence-first eval infrastructure for AI-powered vuln find/validate/patch — maps to shipping harnesses, metrics skepticism, agent measurement. Big-company stability + safety-adjacent cyber.

Requirements (public): Bachelor’s or equivalent + 5+ years ML eng/evaluation; designing benchmarks/metrics; solid Python for shared infra; experiment tracking/data pipelines. Preferred: security evaluation, agent/LLM behavior measurement, public benchmarks.

Logistics: US remote in listed locations (mirrors commonly: Santa Clara + remote CA/NC/NY/TN/FL). Texas not consistently listed on this exact JD — verify on jobs.nvidia.com before prioritizing.

Travel: Unknown; typical corp remote.

Pay: Posted base: $152,000–$241,500. ESTIMATE total comp: NVIDIA ML Eng Levels.fyi US medians often ~$200k–$330k+ TC depending on level (stock-heavy) — not this JD’s offer.

Day-to-day: Build benchmarking/reproducibility systems; define metrics/protocols; map every result to code+runs; keep conclusions reviewable.

Link: NVIDIA job 893396714449

Caveats: State eligibility may block Tyler TX; applications may already be past “accept until July 30, 2026” language on some mirrors — confirm still open. Impact is AI-for-security tooling more than frontier alignment theory.

#5 · Researcher / Senior Researcher · Epoch AI

fully remote benchmarks / measurement

Why for Eric: Fully remote with PT–CET hiring; measurement/critique of AI progress fits “make numbers mean something.” Lower day-to-day red-team than #1–#3, but family-logistics excellent and credibility high.

Requirements: Open to varied backgrounds; researcher roles across multiple teams (incl. Benchmarking Reviews producing critiques of AI benchmarks). Prefer overlap with PT and UTC; can travel to ~3 retreats/year.

Logistics: Fully remote; many countries PT→CET. America/Chicago OK.

Travel: Prefer candidates who can attend ~3 staff retreats/year.

Pay: Posted examples: Researcher/Senior across teams often cited in wide bands (e.g. Benchmarking Reviews mirrors $100k–$200k); broader researcher postings sometimes higher — confirm per team. Do not assume top of range.

Day-to-day: Research trends/capabilities; write reviews/critiques; publish for policymakers and industry.

Links: Epoch Lever researcher · epoch.ai careers

Caveats: More research writing than agent harness shipping; retreat travel; not adversarial red-team primary.

#6 · Research Engineer, Model Evaluations · Anthropic

remote-friendly + travel 25%+ office exceptional pay

Why for Eric: Design/run capability + safety evals; harden distributed eval platforms against training checkpoints; dashboards under time pressure — almost a job description of “ship eval harnesses.” Lab-scale impact.

Requirements: Strong Python; distributed systems/data pipelines; clear communication; on-call comfort during training runs. Preferred: LLM scaffolding, dashboards, eval metrics, stats, ML infra.

Logistics: SF or NYC; hybrid ≥25% in office; labeled Remote-Friendly (Travel-Required).

Travel: Material — plan regular SF/NY presence from Texas.

Pay: Posted: $500,000–$850,000 annual salary (Greenhouse).

Day-to-day: New evals (reasoning, agents, safety); distributed eval execution; mid-run debugging; researcher partnership.

Link: Greenhouse 5198255008

Caveats: Family travel cost is the real filter. Extremely competitive. Visa sponsorship possible but not guaranteed for every candidate.

#7 · Red Team Engineer, Safeguards · Anthropic

remote-friendly + travel product red-team

Why for Eric: Jailbreaks, agentic prompt-injection, automated testing frameworks — closest lab red-team title to bounty harness work.

Requirements: Pentest/red-team/appsec experience; model jailbreaking; agentic workflow testing; AI/ML security preferred; API/authz abuse, T&S backgrounds helpful.

Logistics: SF; Remote-Friendly (Travel Required); ≥25% office policy.

Travel: Yes — hybrid expectation.

Pay: Not clearly listed on the public card fetched todayunknown (ask). Sibling Anthropic RE roles post very high bands; do not invent a number.

Day-to-day: Creative multi-step attacks on product surfaces; systematic methodologies; automated LLM testing frameworks.

Link: Greenhouse 5320469008

Caveats: Hybrid travel; may prefer classical security pedigrees alongside LLM jailbreak skill. Competition high.

#8 · Engineering Manager, Red Team · FAR.AI

remote manager

Why for Eric: If multi-agent fleet coordination / Board-style review generalizes to leading a red-team eng org — remote path into FAR red-team. Skip if you want IC-only.

Requirements: Mid experience managing eng delivery (confirm Ashby JD).

Logistics / travel: Remote; travel unknown.

Pay: Posted (80k Hours): $170–$250k.

Day-to-day: People/project systems enabling red-team research velocity.

Link: Ashby EM Red Team

Caveats: Manager ≠ IC red-teamer; hiring committees weight management receipts.

#9 · Research Engineer, Frontier AI Risk Management · SaferAI

remote considered EU preference

Why for Eric: Translate risk commitments into concrete evals/mitigations/red-teams of mitigations — eng-adjacent assurance work.

Requirements: Technical RE skills for evals/mitigations; async remote collaboration; strong writing.

Logistics: Team mostly Paris/London; remote for particularly strong candidates.

Travel / pay: Unknown. Pay: unknown (not posted on page fetched).

Link: safer-ai.org jobs (FAIRM Research Engineer) — verify live; some mirrors 404’d on fetch.

Caveats: EU timezone bias; page may close when filled; weaker US remote guarantee than METR/FAR/Epoch.

#10 · MTS, Safety for Agents · Cohere

remote-friendly ~50% UK/EU overlap

Why for Eric: Safety for agents that take actions — data gen, post-training, evals. Multi-agent shipping experience is relevant.

Requirements: Strong SE + stats/experimental design; data collection with annotators; Python/ML frameworks.

Logistics: Remote-friendly across NA/Europe/etc.; ~50% working-day overlap with UK/EU (US East fine; Chicago is early mornings).

Pay: Third-party mirrors have cited CAD bands — treat as ESTIMATE / unverified here; confirm on Cohere careers.

Day-to-day: Safety post-training + eval methods for tool-using models; cross-functional with product/policy.

Caveats: Timezone tax from Texas; confirm role still open.

Honorable mentions (logistics hard)

  • Apollo Research — AI Red Team Engineer: Superb scheming/control/monitor red-team fit (Watcher). In-person London or SF with flexible WFH. Posted SF $182k–$238k / London £122–160k. Lever
  • Redwood Research — Member of Technical Staff: AI control, alignment-faking, safety cases. Berkeley onsite. Posted $350k–$850k. careers
  • Anthropic — Cyber Evaluations Engineer: Cyber capability + safeguard robustness evals; SF | DC; hybrid policy. Greenhouse
  • OpenAI Preparedness (Automated Red Teaming, Threat Modeler, etc.): high fit, San Francisco. ART example posted ~$295k–$445k.
  • xAI / SpaceX AI safety eng: Palo Alto / office-heavy (e.g. Imagine Safety; Safety Security).
  • METR MTS (Eval Execution, Embedded Assessments, Security Eng, Cyberforensics): Berkeley onsite/hybrid; posted roughly $328k–$687k depending on seat.
  • NIST CAISI: US citizenship; in-person DC & SF; Research Engineer/Scientist and MTS hiring — not remote-first. nist.gov/caisi
  • UK AISI: Strong red-team/control openings (£65–145k bands) — London / UK, visa. Closes noted on some RE roles ~30 Sep 2026.

On-ramps / not FTE jobs

  • BlueDot Impact — Technical AI Safety (free course): bluedot.org
  • MATS research fellowships: matsprogram.org
  • ARENA technical bootcamps: arena.education
  • Anthropic Fellows (AI Safety / AI Security): 4 months FT; remote-friendly US/UK/Canada with work auth; stipend $3,850/week USD (+ ~$15k/mo compute). On-ramp with historically strong conversion to Anthropic safety FTE. Fellows Greenhouse
  • Anthropic model safety bug bounty + red.anthropic.com — already enrolled; finish a clean claim.
  • xAI/SpaceX HackerOne (authorized only): hackerone.com/x
  • 80,000 Hours map + board: 80000hours.org/ai · jobs.80000hours.org
  • METR General Expression of Interest (flexible): metr.org/careers
  • Hawk open eval stack: hawk.metr.org — practice Inspect-like task work publicly.

Suggested apply order (family-aware)

  1. METR Task Development Engineer + FAR RE / Jailbreaking Lead (true remote).
  2. Epoch researcher + NVIDIA eval eng if TX (or relocate-state) eligibility clears.
  3. Anthropic Fellows (time-boxed on-ramp) or Anthropic Model Evals / Safeguards Red Team if hybrid math works.
  4. Keep Apollo/Redwood/OpenAI/xAI/CAISI as stretch only if relocation or rare-travel exceptions appear.

Parallel portfolio (1–2 weeks): one authorized bounty submission; one public reproduction of a METR/Inspect-style eval with a clean negative-or-partial result write-up; link Hawk familiarity.

Sources & honesty

Checked 2026-09-15 (UTC) via WebFetch/WebSearch/curl: metr.org/careers + Lever; far.ai / Ashby / jobs.80000hours.org; Anthropic Greenhouse; NVIDIA job page mirrors; Epoch Lever; Apollo Lever + careers FAQ; Redwood careers; NIST CAISI; UK AISI EOIs; OpenAI/xAI search results; article on-ramps in /workspace/eric-jobs/labs.txt & long.txt.

Machine-readable summary: /workspace/eric-jobs/ranked-jobs.json.

No secrets. No fabricated private résumé bullets. If a posting closed between research and your click, trust the live page.