Data Scientist & Data Analyst · M.S. Data Analytics Engineering @ Northeastern
I build end-to-end ML and analytics that actually ship — a win-probability model that grades every NFL fourth-down call, and a pipeline that has tracked multi-cloud GPU prices every week since June.
Northeastern University
M.S. Data Analytics Engineering
Expected May 2027 · Boston, MA
GPA: 3.6/4.0
National Changhua University of Education
Bachelor of Business Administration
January 2025 · Changhua, Taiwan
I am a graduate student in Data Analytics Engineering at Northeastern University, passionate about turning complex data into clear, actionable business insights.
I specialize in machine learning, data visualization, and cross-functional storytelling to drive real-world business impact.
Most A/B projects run one t-test on a downloaded CSV. I built the whole platform — assignment, sequential and Bayesian tests, CUPED, SRM — and then proved it works. Across 1,000 simulated experiments with a known answer, A/A false positives land at 5.3% against a nominal 5%; checking the results ten times inflates an ordinary test to 19.5% while mine holds 0.9%.
Ask a question in plain English; Claude writes the SQL and DuckDB runs it inside your browser against a 1.16M-row, 13-table warehouse — no server. I benchmarked the models on the same 13 questions with an independent model as the judge: Opus 69%, Haiku 38%. That gap is why the app generates SQL with Opus and only lets Haiku summarise the answer.
An ML system that predicts MLB games and 2026 playoff odds — five win-probability models, 10,000 Monte Carlo season simulations, retrained and republished automatically every morning, with every prediction logged before first pitch.
A pipeline that has been collecting every week since June — 47,123 GPU price points across all three clouds so far, plus ten years of market share rebuilt from SEC filings. Five live pages, a SQLite star schema, and an event study on whether ChatGPT bent the revenue curve.
Mapped 16 years of U.S. aerospace patents (5,200+) across 300+ metro areas — and caught a silent data bug that had erased Los Angeles, the #1 metro, from the map; fixed it with CBSA-code joins and validated against official USPTO totals.
Graded all 15,545 NFL fourth downs from 2021–2024 against a win-probability model I fitted myself. It agrees with the coach 65% of the time; the disagreements are worth 22 wins a season league-wide. Ships with a calculator that answers any fourth down instantly.
The crime file I was scoring on had exactly 16,000 rows — and the county API that produced it returns at most 16,000 rows per request. I rebuilt it from source into 25,829 incidents, which changed the ranking, then measured how much of that ranking was ever real: across 10,000 weightings the median area moves 8 places out of 17.
Featured