Baseten · Research profile

Validate post-training against production behavior.

Mohamed A M Elansary, PhD — scientific ML, measurement under uncertainty, hydroclimate forecast evaluation, and production agent evaluation sets.

Post-training validation mindsetInference-adjacent measurementUncertainty quantificationProduction agent evals

Scientific measurement

  • Designed multi-model forecast comparisons across basins and hydroclimates.
  • Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
  • Ran reproducible Python, R, Bash, Linux, and HPC workflows.

Production systems

  • Builds GPT, Claude, and Gemini agent workflows at Vertexium.
  • Maintains regression evaluation sets for production agent behavior.
  • Ships retrieval, routing, tenant isolation, provenance, and validation systems.

Proposed measurement approach

Define intended behavior for a post-training or inference-adjacent question; build a small evaluation set with provenance and ambiguity labels; implement statistical analysis, stratification, and uncertainty; compare simple baselines; and report what the signal does and does not support before expanding into larger training or serving loops.

Honest fit boundary

This seat has less explicit benchmark and evaluation ownership than dedicated evals roles. Baseten-scale inference serving, multi-node trillion-parameter training leadership, RLHF, and alignment research are a stretch. I do not invent post-training metrics. My contribution is scientific evaluation under uncertainty, HPC rigor, and production agent regression evaluation.

Role and location

Post-Training Research Scientist · Ashby fields: Location San Francisco · Location Type Hybrid. Relocation with a support package is an honest discussion point; remote eligibility is not asserted.

$210K – $285K · Offers Equity · Official role posting