Home About Skills Projects Experience Education Contact
Data Science Graduate

Hi, I'm HASNAIN SIZAR

Building end-to-end data systems

UC Irvine Data Science graduate building analytics tools end to end in Python and SQL, from API ingestion and relational storage to detection rules validated against labeled ground truth. Every project ships with tests, documentation, and CI.

Connect
Illustrated portrait of Hasnain Sizar
Illustrated portrait of Hasnain Sizar

Data Scientist

Analytics pipelines, machine learning and statistical systems, revenue analytics tools.

Chino, CA · Open to data science roles

I'm Hasnain Sizar

I graduated from UC Irvine with a B.S. in Data Science in June 2026. My projects run end to end: ingesting public APIs and flat files into relational storage, engineering features and detection rules, and validating the output against labeled ground truth instead of eyeballing it.

Python and SQL are the core of my work, with machine learning and statistics on top: supervised models, probability modeling, A/B testing, and causal inference with propensity score matching and difference-in-differences.

I care about code other people can run. Everything I ship has tests, documentation, and CI, and states the tradeoff behind every threshold. Two of my projects are built for revenue teams, surfacing expansion and competitive displacement signals and mapping each one to a follow-up action.

5 End-to-End Projects
266+ Tests on Green CI
68,404 Matches in Causal Panel

Skills & Tools

The languages, methods, and tooling behind the projects below.

Languages

Daily drivers for analysis and pipelines

Python SQL R

Python Ecosystem

Modeling and analysis libraries

Pandas NumPy Scikit-learn Statsmodels Matplotlib

Databases & Data Engineering

Ingestion, storage, and data quality

PostgreSQL SQLite Schema Design Window Functions CTEs ETL Pipelines REST API Ingestion Web Scraping Selenium Deduplication Data Quality Checks

Machine Learning

Supervised models, evaluated honestly

Supervised Classification Regression Feature Engineering Model Selection Model Evaluation Precision Recall AUC Cross-Validation Probability Modeling

Statistics & Experimentation

Testing, inference, and causal methods

A/B Testing Hypothesis Testing Regression Time Series Causal Inference Propensity Score Matching Difference-in-Differences Event Study Placebo Tests

BI & Visualization

Dashboards and automated reporting

Power BI DAX Power Query Tableau Excel PivotTables Automated HTML Reporting Jinja2

GTM & Revenue Systems

Signals and workflows for revenue teams

Clay n8n Enrichment Data Hygiene Lead Scoring Routing Logic Product Usage Signals Expansion Signal Detection Competitive Displacement Signals Outbound Personalization REST API Automation

Tools & Practices

How the code gets shipped

Git GitHub GitHub Actions CI pytest ruff mypy Jupyter Typer Unix Command Line Agile/Scrum

Selected Projects

Five end-to-end systems. Each card shows real output, and each repo ships with tests and documented tradeoffs.

Groundswell expansion brief for a sample account showing a composite expansion score of 37 and a consumption ramp signal Analytics

Groundswell

Account usage analytics and alerting over daily telemetry (compute, seats, workspaces, feature breadth) for a 60-account portfolio. Five rule-based detectors using window medians and weekday-matched baselines separate durable change from noise, feeding composite scoring and automated HTML briefs. Validated at 0.89 precision and 0.94 recall over 10 generated datasets and 600 accounts. 116 tests, green CI.

Python SQLite Typer Jinja2 GitHub Actions
Snowwatch displacement digest showing 43 signals collected over 14 days and a scored displacement signal ETL + Scoring

Snowwatch

Competitive-intelligence signal pipeline that ingests posts and job listings from three public APIs, deduplicates into SQLite, and applies direction-aware rules to surface platform-switching signals mapped to follow-up actions. False positives cut with a skills-list dampener, staffing-firm flags, score floors, and suppression notes. 150 tests, green CI.

Python SQLite Typer Hacker News API Stack Exchange API Adzuna API
Rxdelta report header showing coverage changes between two monthly CMS Part D snapshots across 5,517 plans CLI + Reporting

Rxdelta

Medicare Part D formulary change monitor. Loads two monthly CMS releases, roughly 1.1M formulary rows per month across 5,518 plans, into partitioned SQLite with a full audit trail, then classifies tier moves, prior authorization, step therapy, quantity limits, additions, and drops and ranks them by estimated member cost impact. Ships documented cost-impact ranges, mypy strict, ruff, a pytest coverage gate, and a self-contained HTML report.

Python SQLite Typer pytest GitHub Actions CMS public-use files
Scatter plot of expected goal difference against final league position for Serie A 2019/20 with a Spearman correlation of 0.987 Causal Inference

Causal Impact of Mid-Season Managerial Changes

Propensity Score Matching combined with Difference-in-Differences to estimate the causal effect of mid-season manager firings across 68,404 matches, the top 20 European leagues, and six seasons (2019/20 to 2024/25). Estimated +0.292 xGD per match over 12 matchweeks, 95% CI [0.192, 0.392], p < 0.001. Pipeline built on API-Football and Transfermarkt via Selenium into a 6-table SQLite database with 2,053 firings, validated with covariate balance, event-study pre-trends, and placebo tests.

Python Scikit-learn Statsmodels Pandas SQLite Selenium
The Oracle prediction card for Norway against England picking England at 64 percent with narration Prediction

The Oracle

World Cup 2026 prediction bot. A transparent Elo-style rating model computes win probabilities and passes only the computed numbers to an LLM for commentary, so the narration cannot invent scores or statistics. Renders shareable 1200x720 PNG cards with team flags, pick, probability split, and narration, plus offline fallbacks, cached assets, and a terminal card for local runs.

Python Anthropic API Pillow

Work Experience

Six years of hospitality operations at a Choice Hotels franchise, from the front desk to supervising it.

Front Desk Supervisor

Sep 2024 to Present

Rodeway Inn (Choice Hotels franchise), Artesia, CA

  • Supervise front desk operations across daily shifts, including agent scheduling, shift handoffs, and cash and credit reconciliation.
  • Train and onboard front desk agents on the property management system, reservation and group booking procedures, and brand service standards.
  • Serve as escalation point for billing disputes and service recovery.
  • Track occupancy, rate, and guest satisfaction results and use trends to adjust staffing coverage and upsell targets.

Front Desk Associate

Jun 2020 to Sep 2024

Rodeway Inn (Choice Hotels franchise), Artesia, CA

  • Managed guest check-in and check-out, reservations, room assignments, and payment processing.
  • Resolved billing corrections, room changes, and service complaints, coordinating with housekeeping and maintenance.
  • Promoted loyalty enrollment and paid room upgrades while meeting property targets.

Education & Certifications

The degree behind the projects, plus the certifications that back the BI work.

Bachelor of Science in Data Science

University of California, Irvine · June 2026

Relevant coursework
Machine Learning Statistical Analysis Data Structures and Algorithms Database Systems Data Mining Linear Algebra Probability and Statistics

Microsoft Certified: Power BI Data Analyst Associate (PL-300)

Microsoft

In Progress

Data Scientist with Python

DataCamp

2023

Let's Connect

Open to data scientist, analyst, and data engineering roles. Email is the fastest way to reach me.

Email

hasnainsizar@outlook.com

Phone

562-386-4852

Location

Chino, CA

Find me online
Copied to clipboard