rishi-shah.ipynb Python 3 · ready
[1]
whoami()
available — open to data analyst roles · US / remote

Rishi Shah

Data Analyst · SQL · BI · Python

I turn messy data into decisions teams can act on.

Phoenix-based data analyst with 2+ years turning fragmented data into dashboards, SQL pipelines, and answers stakeholders actually use — with the Python, ML, and NLP chops to push past standard reporting when a problem calls for it.

📍 Phoenix, AZ · open to relocating anywhere in the US
[2]
rishi.summary()

role     = data analyst
focus    = SQL · BI dashboards · ETL · stakeholder reporting
degree   = M.S. Information Technology, ASU · GPA 4.0
experience = 2+ yrs turning fragmented data into dashboards & decisions
edge     = Python, ML & NLP projects that go beyond standard reporting
status   = seeking data analyst & analytics-engineering roles

[3]
rishi.skills.groupby("domain")
#domainstackfocus
0pythonpandas · numpy · scikit-learn · pytorch · xgboostmodeling & analysis
1cloud + mlopssagemaker · lambda · emr · s3 · athena · cloudwatch · snsdeploy & automate
2sql + datasql · mongodb · snowflake · etl · stored procsmodel & query
3nlp + aibiobert · regex ner · rag · transfer learningextract meaning
4viz + bitableau · power bi · matplotlib · seaborncommunicate
5dev toolsjavascript · angular · node · git · c++build the scaffolding
[6 rows × 4 columns]
[4]
rishi.projects.sort_values("impact", ascending=False)
# 01nlp · biobert · pytorch

PubMed NLP Hybrid BioBERT Pipeline

Fine-tuned BioBERT on a Tesla T4 and bolted on a regex layer to catch what the model missed — pulling sugar-sweetened-beverage risk thresholds out of 167 PubMed abstracts. Recall climbed from basically zero to 77%, surfacing 28 high-confidence dose–response findings as clean, structured CSV.

BioBERTPyTorchRegex NERPubMed
# repo: private
# 02aws · xgboost · mlops

Intelligent Heart Attack Prediction System

A five-service AWS pipeline — S3 → EMR/PySpark → SageMaker → Lambda → SNS — that ingests wearable IoT data, scores cardiac risk with a tuned XGBoost model, and fires real-time alerts. Built in a schema-alignment guard so training and inference can't silently drift apart.

SageMakerEMRLambdaXGBoostPySpark
# repo: private
# 03sql · database engineering

Healthcare Database Management System

A fully normalized 3NF schema across seven tables that held zero integrity violations over 100+ records — with an audit trigger logging every change, three analytical views, stored procedures, a scalar UDF, and a cursor-driven glucose-alert routine.

SQL3NFTriggersStored ProcsUDF
view repository
# 04sql · tableau

End-to-End Sales Data Analysis

Pulled and reshaped relational e-commerce data with heavy SQL — joins, CTEs, window functions — then built a Tableau dashboard on trends, regions, and product margin that pointed to a ~10% cross-sell opportunity.

SQLTableauCTEsWindow Funcs
view repository
# 05aws sagemaker · xgboost

Netflix Ratings & Viewing Trends Prediction

Engineered 62 features across 6,187 records and deployed dual XGBoost models to a live SageMaker endpoint — then used CloudWatch to catch a false-positive bias and document the threshold-tuning fix instead of burying it.

SageMakerCloudWatchXGBoostFeature Eng.
# repo: private
# 06python · eda · visualization

IPL Cricket Analytics Dashboard

Ten-plus seasons of IPL data explored with pandas, matplotlib, and seaborn to map player form, team strategy, and what actually swings a match.

PythonPandasMatplotlibSeabornEDA
# repo: private
[5]
rishi.background(order="recent")
Aug 2024 – May 2026
Tempe, AZ
GPA 4.0 / 4.0

M.S. Information Technology

Arizona State University
  • Big data analytics, NLP, data visualization, advanced databases, and system architecture.
  • Every major project shipped as something runnable — AWS pipelines, a BioBERT model, a production-grade SQL database — not slideware.
May 2022 – Jul 2024

Data Analyst

Shiv Infotech
  • Replaced scattered spreadsheet reporting with Tableau and Power BI dashboards that gave 10+ stakeholders one source of truth — cutting time-to-spot-a-bottleneck by roughly .
  • Built Python data-processing pipelines with Pandas and scikit-learn to clean and parse 500K+ record datasets, engineering target features that sharpened internal business-forecasting accuracy — while automating 15+ hours a month of manual cleanup.
  • Rebuilt SQL ELT extraction across 3–5 concurrent client projects, trimming 2 days off reporting turnaround and freeing capacity for more accounts per sprint.
  • Kept Python and Snowflake ETL pipelines accurate and ready for both batch and real-time reporting.
PythonPandasscikit-learnSQLSnowflakeETLTableauPower BI
Oct 2021 – Apr 2022

Software Developer Intern

IT Path Solutions
  • Engineered high-concurrency web applications in Angular and Node.js, adding server-side rendering and code-splitting that boosted performance ~35% and supported a 20% increase in user retention.
  • Shipped a full-stack client-management feature and Tableau dashboards on user-interaction data, lifting task-completion efficiency ~15%.
  • Ran delivery in Agile/JIRA, translated stakeholder requirements into specs, and tracked sprint metrics in Power BI to keep releases on schedule.
AngularNode.jsSSRTableauPower BIAgile/JIRA
[6]
# the longer version
print(rishi.about)

I'm a data analyst who likes to get technical — the kind who learned early that a dashboard nobody trusts is worse than no dashboard at all. Most of my work is turning fragmented data into SQL pipelines, BI dashboards, and answers that a stakeholder can act on Monday morning.

What sets me apart from a typical analyst is how far I'll push when a problem calls for it: a BioBERT NLP pipeline pulling structured findings out of biomedical literature, a five-service AWS system scoring health risk in near real-time, a normalized SQL database that holds its integrity under load. I reach for Python, ML, and cloud tooling when standard reporting runs out of road — but the goal is always the decision, not the model.

I just finished my M.S. in Information Technology at Arizona State (4.0 GPA) and I'm looking for my next role — data analyst or analytics engineer — in Phoenix, remote, or anywhere in the US. If you've got messy data and a decision riding on it, I'd like to talk.

[7]