PORTFOLIO ARNAV RAUT · 2026
ARNAV RAUT

DATA SCIENTIST

B.Tech in Computer Science Engineering @ MIT-ADT University.
Building predictive models, natural language pipelines, and neural recommendation engines.
BACKGROUND CORE DOMAINS & TOOLSTACK
01

What I Explore & Build.

I am a Computer Science Engineering student focused on extracting actionable intelligence from complex, high-dimensional data. My primary interests lie in predictive modeling, deep learning for time series, natural language processing, and bridging the gap between raw research algorithms and practical interactive applications.

Currently maintaining a CGPA of 8.47 at MIT-ADT University, active in technical student chapters, and building production-ready architectures that deliver tangible machine intelligence.

Languages
Python C++ JavaScript R SQL
Core Domains
Machine Learning Deep Learning NLP Time-Series Data Visualization
Current Toolstack
PyTorch TensorFlow scikit-learn Hugging Face BERT Pandas NumPy Streamlit
TIME-SERIES FORECASTING JULY 2025 · TSLA HISTORICAL DATA
02

ARIMA vs LSTM — Stock Price Prediction

Statistical Econometrics vs Deep Sequential Neural Networks

Evaluated whether classical econometric formulations (ARIMA) or recurrent neural architectures (LSTM) better capture volatile price shifts and non-linear trend dynamics. Fitted models against multi-year historical Tesla (TSLA) market trajectories via the Yahoo Finance API with sliding lookback windows and ADF stationarity tests.

PYTHON PYTORCH STATSMODELS STREAMLIT YAHOO FINANCE API
Evaluated Metric LSTM ACHIEVED LOWER RMSE THAN ARIMA
Key Takeaway

While ARIMA efficiently models stationary linear trends, multi-layer LSTMs with dropout significantly outperformed on multi-step horizon forecasts during high-volatility regimes.

NLP & CLASSIFICATION JUNE 2025 · KAGGLE TWITTER CORPUS
03

Understanding the Voice of Twitter

Multi-Stage Feature Extraction & Logistic Classification

Engineered a complete natural language processing pipeline to extract sentiment polarity from informal, noisy social media text. Implemented end-to-end preprocessing: URL/mention scrubbing, tokenization, POS tagging, WordNet lemmatization, and sub-linear TF-IDF representation across 5,000 unigram and bigram features.

PYTHON SCIKIT-LEARN NLTK TF-IDF NLP PIPELINE
Test Evaluation ~70% TEST ACCURACY
Pipeline Architecture

Raw Tweet Text → Regex Normalization → NLTK Lemmatizer → TF-IDF Vectorizer (5k features) → L2-Regularized Logistic Regression Classifier.

NEURAL RECOMMENDERS MAR–MAY 2025 · TRANSFORMER EMBEDDINGS
04

Semantic Similarity with BERT

Content-Based Recommender Using Dense Transformer Embeddings

Overcame vocabulary mismatch and keyword sparsity in recommendation catalogs by synthesizing multidimensional game metadata (titles, user reviews, tags, plot summaries, genres) into rich semantic embeddings via fine-tuned BERT representations, paired with vector cosine similarity ranking.

PYTHON HUGGING FACE BERT COSINE SIMILARITY STREAMLIT
Engagement Metric ~60% HIT RATIO (TOP-K)
Semantic Advantages

Captures contextual synonyms and aesthetic vibes across disparate game genres that traditional TF-IDF or collaborative filtering fail to connect.

OUTLOOK & CONTACT AVAILABLE FOR OPPORTUNITIES
05

Let's Build Something Intelligent.

I am always eager to collaborate on challenging machine learning engineering, data science research, time-series forecasting, or NLP systems.

Whether you are looking for an ambitious intern, project collaborator, or just want to chat about AI architectures, drop a message!

LOCATION & AVAILABILITY
Pune, Maharashtra, India · 2026