Skip to content
IIIT Ranchi • 3rd Year CSE

Ashutosh Pandey

Backend Engineer building AI-powered systems and automation tools.

I build backend systems, AI-integrated applications, automation pipelines, and developer-focused products that solve real-world problems.

Ashutosh Pandey
About

Who I am

I'm a Computer Science Engineering student at Indian Institute of Information Technology, Ranchi, focused on backend engineering and AI-integrated software development.

I enjoy understanding how systems work behind the scenes — from APIs and authentication to databases, automation pipelines, data processing, and LLM-powered applications.

I'm particularly interested in combining traditional backend engineering with modern AI capabilities to build practical software that solves repetitive and complex real-world problems — not demos, but systems people actually rely on.

Engineering Focus

  • Backend Engineering

    FastAPI · REST APIs · Authentication · RBAC

  • AI Integration

    LLMs · LangChain · LangGraph · RAG

  • Automation

    Python Automation · Document Processing · Data Extraction · Validation Pipelines

  • Data & Databases

    PostgreSQL · Supabase · MongoDB · Redis

  • Developer Tools

    Git · GitHub · Docker · Postman

Education

Academic background

Indian Institute of Information Technology, Ranchi

3rd Year8.00 CGPA

Bachelor of Technology — Computer Science and Engineering

2024 – 2028

Coursework and self-directed projects centered on backend systems, databases, and AI-integrated software.

Class XII (Senior Secondary)

Class XII81%

Higher Secondary Education

Completed

Class X (Secondary)

Class X86%

Secondary Education

Completed

Currently

Current Engineering Focus

What I spend most of my time building with right now.

Backend Engineering

Designing REST APIs and service logic that stay maintainable as requirements change.

FastAPIREST APIsAuthenticationRBACAsync ProgrammingAPI Architecture

AI Integration

Wiring LLMs into real application logic instead of treating them as a black box.

LLMsLangChainLangGraphRAGHugging FaceAI Workflows

Automation

Turning repetitive manual processes into reliable, reusable pipelines.

Python AutomationDocument ProcessingData ExtractionValidation Pipelines

Data & Databases

Modeling data and querying it efficiently, sync or async.

PostgreSQLSupabaseMongoDBRedisMulti-Tenant Architecture

Developer Tools

The everyday toolchain behind shipping backend work.

GitGitHubDockerPostmanRenderVS Code
Skills

Technical skills

Languages

PythonC++GoCSQL

Backend

FastAPIREST APIsPydanticSQLAlchemyasyncpgJWT AuthenticationRBACAsync Programming

AI / GenAI

LangChainLangGraphRAGHugging FaceLLM IntegrationPrompt EngineeringAI Workflows

Databases

PostgreSQLSupabaseMySQLMongoDBRedisMulti-Tenant Architecture

Data / ML

NumPyPandasScikit-learnEDAFeature EngineeringMatplotlibSeabornETLOCR

Web Scraping

BeautifulSoupRequestsWeb Automation

Core CS

Data Structures & AlgorithmsDBMSObject-Oriented ProgrammingOperating Systems

APIs & Integrations

Google Drive APIGmail APIOAuth 2.0Google Apps Script

Tools

GitGitHubDockerPostmanRenderVS CodeJupyterGoogle ColabGoogle Workspace
Engineering Knowledge

How a request becomes a response

Backend concepts I understand and apply — hover a stage to see what happens there.

Browser / mobile app sends a request

Identity & Access

AuthenticationAuthorizationJWTRBACOAuth 2.0SessionsRefresh Tokens

API Design

REST APIsAPI ValidationCRUDAsync Programming

Data Layer

Database DesignPostgreSQLMongoDBMulti-Tenant Architecture

Shipping Software

DockerDeployment
Experience

Where I've worked

Engineering decisions and problems solved, not just a list of technologies.

AI/Automation Software Developer Intern · Ambition Colonisers Pvt Ltd

Onsite — Gurugram, Haryana

Aug 2026 – Sep 2026

Built a multi-tenant financial reconciliation platform with schema-per-tenant data isolation, role-based access control, and an automated bank statement ingestion pipeline.

Problem

Reconciling financial data across multiple companies required strict data isolation between tenants and fine-grained permission boundaries, while bank statement handling was manual, format-inconsistent (PDF/XLSX/XLS/CSV), and didn't scale as more companies were onboarded.

Approach

Built a schema-per-tenant PostgreSQL architecture behind a FastAPI backend, layered RBAC with JWT authentication enforced at both the API and UI layers, and automated statement ingestion — including a Gmail-to-Drive collection step — with async background processing so large imports never timed out.

  • Built a multi-tenant financial reconciliation platform (Python, FastAPI, PostgreSQL, React 18, Vite, Tailwind CSS) serving 8 isolated company schemas with schema-per-tenant data isolation
  • Designed and implemented an RBAC system with JWT authentication and 4 permission tiers, enforced at both API and UI layers
  • Built an automated bank statement ingestion pipeline supporting PDF, XLSX, XLS, and CSV formats, parsing 300+ statements with configurable page ranges and batch processing
  • Engineered a Gmail-to-Google Drive automation using Google Apps Script and the Gmail/Drive APIs with OAuth 2.0, removing manual statement downloads
  • Implemented asynchronous background job processing with real-time progress tracking and polling-based status updates to prevent timeouts on long-running imports
  • Maintained 85+ automated integration tests validating business rules, permissions, and data integrity across tenant schemas
PythonFastAPIPostgreSQLReact 18ViteTailwind CSSJWTRBACGoogle Apps ScriptOAuth 2.0

GenAI Intern · MySSCGuide

Remote

Jan 2026 – Apr 2026

Built a data engineering pipeline that turns messy educational content from many sources into structured, validated JSON datasets.

Problem

SSC CGL/CHSL exam question data existed across 10+ scattered sources and formats — scanned PDFs, inconsistent layouts, bilingual (English–Hindi) content — with no structured, machine-readable form.

Approach

Combined web scraping, OCR-based document processing, and an ETL pipeline to extract, normalize, and validate the data, using an LLM with custom validation to assist with structuring question data at scale rather than relying on brittle manual rules alone.

  • Engineered Python automation architectures that scaled content generation, reducing manual workload by 40% and saving 15+ hours weekly
  • Built Python ETL pipelines ingesting data from 10+ web sources, processing 50K+ raw educational records into normalized formats
  • Integrated LLMs with custom validation to generate structured educational content in JSON, reducing API schema errors by 30%
  • Automated data extraction and AI-driven content workflows via OCR, improving data accuracy by 60% and generation speed by 40%
  • Partnered with a 5-member team using Git/GitHub to resolve merge conflicts and integrate 10+ features
PythonOCRETLWeb ScrapingRegexJSONLLM Integration
Projects

What I've built

Each one expands into the problem, the engineering decisions, and the tech behind it.

Problem Solving

Problems I've solved

Not tutorial projects — real constraints, real trade-offs.

Problem Solving & DSA

I practice data structures, algorithms, and competitive problem-solving in C++ — 300+ problems solved across LeetCode, GeeksforGeeks, and CodeChef with a focus on Medium/Hard difficulty, plus 30+ CodeChef contests.

Data Structures & AlgorithmsBinary SearchStackQueueLinked ListMonotonic StackMatrix BFSProblem Solving

Problem

Financial reconciliation across multiple companies needed strict data isolation and role-based permission boundaries, with manual statement handling that didn't scale.

Approach

Schema-per-tenant PostgreSQL architecture + RBAC with JWT + an automated multi-format statement ingestion pipeline.

Outcome

A platform serving 8 isolated company schemas, parsing 300+ statements automatically and covered by 85+ integration tests validating permissions and data integrity.

System: Multi-Tenant Financial Reconciliation Platform

Problem

Large volumes of educational data scattered across 10+ inconsistent sources.

Approach

Web scraping + ETL pipeline + normalization + LLM-assisted validation.

Outcome

Structured, validated datasets (50K+ records) suitable for downstream applications, cutting manual workload by 40%.

System: Educational Data Pipeline

Problem

Processing scanned, bilingual educational PDFs into structured question data.

Approach

OCR + extraction + LLM processing + validation + duplicate filtering.

Outcome

Machine-readable, deduplicated question datasets, with data accuracy improved by 60% and generation speed by 40%.

System: Educational Data Pipeline

Resume

Get the full picture

A structured, one-page summary of my experience, projects, and skills.

Contact

Have a problem worth solving?

I'm interested in backend engineering, AI-powered systems, automation, and building practical software.