Measure physical-world generalisation.
Mohamed A M Elansary, PhD — scientific ML, measurement under uncertainty, hydroclimate forecast evaluation, and production agent evaluation sets.
Scientific measurement
- Designed multi-model forecast comparisons across basins and hydroclimates.
- Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Python, R, Bash, Linux, and HPC workflows.
Production systems
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Maintains regression evaluation sets for production agent behavior.
- Ships retrieval, routing, tenant isolation, provenance, and validation systems.
Proposed measurement approach
Define intended behavior for a long-horizon prediction, scene-fidelity, or driving-performance question; build a small evaluation set with provenance and ambiguity labels; implement statistical analysis, stratification, and uncertainty; compare simple baselines; and report what the signal does and does not support before expanding into simulator or policy loops.
Honest fit boundary
This role is less direct evaluation work than dedicated evals seats. Autonomous-driving and robotics reinforcement learning, world-model training, sim-to-real robotics transfer, and RLHF are a stretch. I do not claim AV safety research or invented driving metrics. My contribution is scientific evaluation under uncertainty, HPC rigor, and production agent regression evaluation.
Role and location
Research Scientist, Wayve Labs · “This is a full-time role based in our office in Sunnyvale based in Sunnyvale, CA (hybrid)”. The posting offers “Relocation support with visa sponsorship”. No clearance, citizenship, or ITAR requirement appears in the posting. Remote eligibility is not asserted.
“ranges from $230,000 to $380,000, plus a competitive equity package” · Official role posting