Welcome

to

Arnav's

Portfolio.

Arnav Prashant Bule

Arnav Prashant Bule

TECH ENTHUSIAST | AI/ML DEVELOPER

Driven tech enthusiast with a knack for problem-solving and a passion for turning ideas into impactful solutions.

  • Passionate about AI/ML, Cloud, and creative tech solutions
  • Blend of coding and creativity: Python + C++ or exploring new AI/ML use-cases
  • Loves building tech with impact and mentoring peers
  • Based in Pune, Maharashtra

Skills

Programming Languages:

PythonC++JavaScriptTypeScript

Data Engineering:

PySparkDelta LakeDatabricksAuto LoaderETL PipelinesMedallion Architecture

AI/ML & Data Science:

MLflowScikit-learnPandasNumPyNLPMLOpsRAGPrompt Engineering

Web Development:

ReactNext.jsNode.jsExpress.jsREST APIsTailwind CSSChrome Extensions

Cloud Computing:

IAMKubernetesVMsDBsDatabricks JobsDABsCI/CDDocker

Data & Databases:

SQLBigQueryMongoDBRedisPostgreSQLLanceDB

Tools & Platforms:

Google Gemini APIOpenAI APIsDeepSeekAWSGoogle Cloud PlatformGitHubMLflowJestOAuth 2.0

Operating Systems:

LinuxWindows

Soft Skills:

Project ManagementCommunicationTeamworkLeadershipCollaborationProblem Solving

Certifications

Databricks Certified Data Engineer Professional
Databricks Certified Data Analyst Associate
Databricks Certified Generative AI Engineer Associate
Databricks Certified Machine Learning Professional
Machine Learning Specialization — Coursera
Google Project Management Professional Certificate
Google Data Analytics Professional Certificate
Networking Basics — Cisco Foundations
DSA to Development — Geeks for Geeks (ongoing)
AWS Certified Solutions Architect – Associate
AWS Certified Machine Learning Engineer – Associate
Cloud Career Practitioner Certified — AWS & GCP

Experience

Associate Data Scientist Intern

Sept 2025 – Present

V4C.ai

  • Built and maintained end-to-end ETL pipelines using PySpark, Delta Lake, and Auto Loader, implementing Medallion Architecture (Bronze–Silver–Gold) for scalable, production-grade analytics and ML workloads.
  • Orchestrated and automated production pipelines using Databricks Jobs and Databricks Asset Bundles (DABs), with parameterization, scheduling, and environment-aware deployments aligned with CI/CD principles.
  • Integrated MLflow for experiment tracking, model versioning, and reproducibility, supporting MLOps workflows and lifecycle management in cloud-based environments.
  • Optimized data transformations, compute usage, and storage layouts to improve performance, reliability, and cost efficiency, enabling downstream BI reporting and ML-ready datasets.

Gen AI Intern — CTO Office

Sept 2025 – Mar 2026

Persistent Systems

  • Designed and developed a full-stack AI email assistant (Email Digital Twin) using a custom Chrome Extension and Node.js/Express backend to automate hyper-personalized email drafting.
  • Integrated Google Gemini (gemini-2.5-flash) to analyze users' sent emails, dynamically extracting communication patterns to build distinct, context-specific writing personas.
  • Implemented Retrieval-Augmented Generation (RAG) with LanceDB vector database, enabling semantic search over past email threads for context-aware draft generation.
  • Engineered a high-performance batch processing system using parallel fetching (10 concurrent requests), accelerating email analysis and persona generation by 5–7×.

AI Research and Development Intern

Jul 2024 – Dec 2024

IS360 Technologies

  • Developed a machine learning pipeline using Linear Discriminant Analysis (LDA) to classify EEG signals from the Auditory Oddball paradigm, detecting event-related potentials like the P300 wave.

Vice President

Jul 2022 – Apr 2024

Cloud Computing Club

  • Led a 55‑member team delivering workshops, seminars, and hands‑on cloud‑training sessions.
  • Grew a community of 600+ students, achieving 90% repeat‑engagement intent.
  • Managed logistics, marketing, speakers, and budgets for large‑scale events, ensuring flawless execution.

Management Associate

Jul 2023 – Present

CodeChef Campus Chapter

  • Managed the chapter's annual calendar and budget, aligning eight coding events per semester with academic schedules.
  • Organised & hosted monthly CodeChef challenge mirrors, boosting average participation from 120 → 350 students.
  • Led cross‑functional sub‑teams (marketing, problem‑setting, tech) and produced run‑books that cut future planning time by 40%.
  • Mentored a 10‑member junior committee through weekly stand‑ups and retrospectives, building a sustainable leadership pipeline.

Projects

Cadence agent playground — a billing tweet classified as billing_or_charge at 0.95 confidence and escalated by rules, beside the drafted reply, the Rules → Retrieve → Gemini → Decide timeline and the cited evidence

Cadence

Sep 2026

Python · FastAPI · Google Gemini · BM25 · Pydantic · SQLite · React 18 · TypeScript · Vite · Tailwind · Vercel

Built an evaluated AI support agent for @SpotifyCares as the Hiver SDE Intern take-home: one structured Gemini call classifies a tweet into 12 intents, drafts a ≤280-character reply grounded in six BM25-retrieved threads from 27,627 real conversations, and decides whether to auto-handle or escalate — with deterministic money/security/legal/churn rules that can force an escalation the model cannot undo, and a release veto that replaces any draft carrying an uncited claim or unsupported link with a holding reply.

Treated the proof as the product: a 250-tweet golden set labelled twice against a written guide and adjudicated (κ 0.96 intent / 0.95 escalation), four baselines, paired-bootstrap 95% CIs, and a blind two-order LLM judge on a different model — the agent's replies beat the nearest historical reply by +1.42/5 (CI 1.15–1.67) across 200 messages, with the judge's 73.5% order consistency and its blind spots published alongside.

Froze a fresh 200-tweet holdout behind SHA-256 locks — sample, labels, threshold and every source file hashed before inference, and a runner that refuses to execute at any other code or label revision — then measured 0.938 escalation recall, 0.789 macro-F1 and 35% auto-handling; replaying the same 800 cached receipts through repaired guards lifted recall to 97.5% with zero new model calls, reported as a retrospective regression rather than fresh evidence.

Made every number reproducible without an API key: all 1,771 Gemini calls live in a committed SQLite replay cache keyed by model, prompts, schema and temperature, so one command recomputes the headline metrics in about 1.4 seconds on Linux and Windows CI — while the client survived the free tier with pooled-key rotation, per-key sliding-window rate limits, 429 cooldowns parsed from server retry hints and deadline-bounded retries.

Deployed on Vercel as a slim Python serverless function (FastAPI, BM25 index rebuilt from the committed corpus at cold start, two-slot concurrency, sanitized errors, no visitor prompts persisted) behind a React dashboard with a live playground, bring-your-own-Gemini-key support held only in sessionStorage, and HMAC-signed HttpOnly admin sessions gating the internal evaluation, golden-set and failure-mode pages.

Parakh results view — extracted questions with score pills and AI feedback beside the handwritten answer sheet, with the selected answer's ink region highlighted on the page

Parakh

Aug – Sep 2026

Next.js 15 · React 19 · TypeScript · Tailwind v4 · pdf.js · Google Gemini · Vercel

Built Parakh (परख), a serverless AI exam checker: upload a question paper and a student's handwritten answer sheet (plus an optional marking scheme) and it extracts every question, finds and highlights each answer's exact ink region on the sheet, grades it, and writes per-question feedback with a teacher summary — no auth, no database, nothing stored beyond the browser session, and a one-click sample exam to try the whole flow.

Kept the backend stateless by rendering documents in the browser with pdf.js — up to 15 pages as 1200px JPEGs with the text layer extracted alongside — so Vercel functions only ever receive compact images; digital papers take a text-only fast path while scanned ones fall back to page-image vision.

Replaced the original local pipeline (DeepSeek-OCR-2 grounding on an RTX 4060 plus GPT-5.6 via the Codex CLI, behind a FastAPI worker, polling job store and Cloudflare tunnel) with two Gemini structured-output calls: one extracts questions with sub-parts split, unprinted marks estimated and OR-choice groups tagged; one multimodal call reads every answer page at once and returns per-answer bounding boxes on a 0–1000 grid, scores and feedback.

Wrote a raw-fetch Gemini client with free-tier key rotation: any number of keys round-robin, advancing on 429/500/503 and stepping down a model ladder before failing loudly, so several free quotas pool into one.

Handled messy scripts in post-processing — boxes clamped and merged when vertically adjacent, scores capped at max marks, skipped OR alternatives shown as "OR — skipped" with each choice-set counted once, stray writing surfaced as "Unmatched answers" — and drew percentage-positioned overlays that track the ink at 50–200% zoom, with invisible hitboxes so clicking an answer on the sheet selects its question.

Sentinel demo on Hugging Face — hero with the #1 hackathon, 67.8-point and 4.96× real-time badges above a gallery of sample clips (fire, smoke, flood, crash) marked Detected · matches ground truth

Sentinel

Sep 2026

Python · PyTorch · CLIP ViT-B/32 · Qwen3-VL-2B + LoRA · YOLO11n · OpenCV · React 19 · Vite

Built a fully local video anomaly detector in a one-day hackathon that runs on a single 8 GB RTX 4060 laptop GPU: a trained head on frozen CLIP ViT-B/32 features scores footage at 2 FPS, a LoRA-adapted Qwen3-VL-2B verifies scene context on a bounded cadence, and per-class temporal state machines turn repeated evidence into timestamped events — zero hosted model calls at runtime.

Placed #1 on the AHC Visual Intelligence Hackathon live leaderboard with 67.8 points, 2.3 clear of second place, processing 47.3 minutes of footage in 9.55 minutes — 4.96× real time including model loading, inference, refinement and explanation.

Trained a 1.61M-parameter rank-8 LoRA on Qwen3-VL-2B in 240 updates, supervising only the 12 answer-letter logits so inference is a single constrained-choice forward pass — no JSON generation and zero parse errors across 932 VLM calls — lifting balanced-holdout macro F1 from 29.5% to 40.3%.

Recovered the hardest long-video class without touching the weights: YOLO11n vehicle proposals, ORB/RANSAC camera-motion compensation and occlusion-tolerant background-relative tracking flagged a truck stationary across 22 samples, which the frozen VLM confirmed as a highway stop — lifting the long-context score from 11.3 to 20.5 with no new false alarms.

Shipped a React 19 + Vite review dashboard over a range-request Python server with click-to-seek event timelines and per-class score histories, plus a leak-proof pipeline — content-hash-grouped holdout, offline Hugging Face cache, and an exporter that rejects inconsistent runtime metadata — covered by 36 passing tests.

QR Studio — a WebGPU voxel tree in the Autumn theme (dark mode) whose golden canopy encodes a scannable QR code for www.arnavbule.in, above the message input and season picker

QR Studio

2026

WebGPU · WGSL · React 19 · Vite · jsQR · Vercel

Built a WebGPU studio that grows any link or message into a 3D voxel tree whose canopy encodes a working QR code — click the tree and it flattens into a top-down view, captured straight off the GPU canvas as a PNG that ordinary phone cameras can scan.

Fixed dark-mode scannability without redesigning the code: the dark theme now decodes reliably on real phone cameras while keeping the tree-derived artwork intact — no black-and-white fallback, no painted frame, and the QR payload, matrix, and finder geometry left untouched.

Shipped four seasonal themes with custom leaf palettes, manually selected light/dark modes with gradual synchronized transitions, seasonal ambient audio, and a responsive desktop/mobile UI — with the chosen leaf color held visually consistent across both themes.

Added QR image decoding by upload, paste, or drag-and-drop with jsQR, so codes can be read back into the studio as well as generated from it.

Tamed a fragile build: the upstream WebGPU engine is pinned to a commit and patched by an ordered chain of source-rewriting scripts, with clean dependency reinstalls enforced on Vercel so cached patched modules can never leak between deployments.

Mini Task Tracker — the Kanban board of a Product Launch workspace with To Do, In Progress, In Review and Completed columns, priority badges and due dates

Mini Task Tracker

2025

Next.js 16 · React 19 · TypeScript · Express · Neon Postgres · Sequelize · Vercel

Built a multi-user task tracker with a Next.js 16 (React 19) frontend and an Express + TypeScript REST API on Neon PostgreSQL via Sequelize: owner-scoped workspaces and tasks (title, description, status, priority, due date, position) shown in Board, List, Table and Timeline views.

Implemented secure auth: 6-digit email OTP verification (10-minute validity) before the first sign-in, bcrypt password hashing, 7-day JWT sessions and password reset via a one-hour emailed link — plus Zod request validation, Helmet headers and per-IP rate limiting (100 requests per 15 minutes).

Built the drag-and-drop Kanban with dnd-kit and persisted every move through one batch endpoint that reorders tasks inside a single database transaction; a freshly verified account is seeded with a Getting Started workspace and sample tasks.

Deployed as a single Vercel project: the compiled Express app is exported as one Vercel Function behind an /api rewrite, with idempotent database initialisation shared by the local server and the function, and Jest + Supertest tests run against a Postgres test database with email delivery mocked.

Email Digital Twin

2025

Chrome Extension · Node.js · Express · Google Gemini · LanceDB · Gmail API · OAuth 2.0

Built an AI-powered Chrome Extension and Node.js backend that learns a user's unique writing style to generate hyper-personalized, context-aware email drafts using Google Gemini and LanceDB.

Integrated Google Gemini's LLM via the API to analyze sent emails, dynamically extracting communication patterns and building distinct writing personas (professional, casual, etc.).

Implemented RAG using LanceDB vector database for semantic search over past email threads, enabling the AI to understand ongoing conversation context for accurate draft generation.

Integrated Gmail API with secure OAuth 2.0 flows within a privacy-first architecture where data processing stays local, and built an injected Gmail UI layer with keyboard shortcuts and a multi-persona editor.

Sicklesense

2025

Python · TensorFlow · Keras · OpenCV · Streamlit · Docker · Hugging Face

Built a low-cost, AI-powered digital telepathology solution that detects sickle cells from blood smear images using a fine-tuned ResNet50 deep learning model, achieving >92% accuracy.

Reduced early screening costs to approximately ₹75 per test — a 95% cost reduction compared to traditional laboratory methods.

Developed an intuitive Streamlit web application providing actionable health guidelines, precautions, and emergency helpline contacts based on AI-generated results.

Enabled telepathology and telemedicine support for frontline health workers in rural areas on low-power edge devices. Secured 3rd place ($1,000) at University of Miami Horizon AI Hackathon 2025 and recognized among Top 100 Startups at IIT Delhi Youth Ideathon.

Road Extraction On Satellite Images

Aug 2024 – Present

SIH Project (No. of Group Members – 6)

Developing software for automated road extraction using CNNs on ISRO's Resourcesat images from the Boonidhi portal.

Built a GUI for specifying areas of interest and generating geographically referenced shapefiles, with email alerts for road changes based on image comparisons.

Optimized for efficient processing of large satellite datasets.

ERP Exerciser

Aug 2024 – Present

Industry Project (No. of Group Members – 3)

Developing ERP Exerciser, a Brain-Computer Interface (BCI) application that leverages Event-Related Potentials (ERPs) to deliver real-time cognitive training, enhancing memory, attention, and executive functions in patients with cognitive impairments.

Prisma

Jan 2023 – Apr 2024

SY Mini Project (No. of Group Members – 3)

Developed an extremely efficient machine-learning application for upscaling and colorizing images, with a model size of just 260 MB and achieving 89% color accuracy in standardized tests (using CNNs and DinkNet).

Potential use cases include military applications, night-vision enhancement, and restoring old photographs.

Technologies and Frameworks I Work With

A comprehensive toolkit for building modern, scalable applications

Python
C++
React
Node.js
Tailwind CSS
MongoDB
MySQL
TensorFlow
OpenCV
NumPy
AWS
Google Cloud
Kubernetes
Docker
Linux
Flutter

Let's Connect and Collaborate!

As an AI and Machine Learning enthusiast, I'm always excited to explore new ideas, share knowledge, and collaborate on projects that push the boundaries of technology. I’m not looking for a role, but I’d love to connect with others who share the same passion for AI, data science, and the endless possibilities they offer.

Interested or want to get to know me a bit better?

Get in Touch

Prefer email?
You can reach out to me at arnav.bule05@gmail.com

Check Out My Resume

Get to Know More About Me

Click below to view or download my latest resume in PDF format from Google Drive.

View Resume

creativity is my craft

ARNAVARNAV