ARYAN.OS — POWER ON
ARYAN.OS
REC
● LIVE  / camera_00  / inference_active

ARYAN
MHALSANK

> Computer Vision Engineer

Building intelligent systems that see, read and reason. From smart glasses for the visually impaired to self-supervised retinal models — I engineer the pipelines between pixels and meaning.

MODELS
24+
PAPERS
04
HACKS
W•5
person 98%
laptop 91%
face 96%
keyboard 87%
monitor 93%
CAM_00 // 1920x1080 // YOLOv8n
INFER
mAP@0.5 : 0.872  |   conf > 0.6
5 OBJECTS
DEPTH MAP
SEGMENT
OCR
> tesseract::detect
> "ARYAN.MHALSANK"
> conf: 0.94 ✓
MOD_02// feature activation across CV / ML / systems

STACK.HEATMAP

lang95
Python
ml88
PyTorch
cv92
OpenCV
cv90
YOLO
cv82
OCR / Tesseract
cv78
Segmentation
ml76
TensorFlow
ai84
LLM / RAG
ai80
Multimodal AI
ml75
Self-Supervised
web80
Django / Flask
web78
React
data74
Supabase / Firebase
ai88
Lyzr Agents
tools70
n8n
sys82
Linux
computer visionmachine learningai / llmlanguage
MOD_03// 6 objects detected in /projects

DETECTIONS.WORKSPACE

6 / 6 LOCKED
INSPECTOR  //  hover an object
> waiting for selection...
MOD_04// experience as training epochs · loss decreasing

TRAINING.PIPELINE

EPOCH_06loss ↓ 0.08
Jan 2026 — Present · Noida, India

Nirikshan AI

AI Intern

Computer vision pipelines — YOLO model fine-tuning, OCR extraction, real-time TTS overlay.

YOLOOCRTTS
EPOCH_05loss ↓ 0.12
Jun 2024 — May 2025 · Kolkata, India

Indian Statistical Institute

Research Intern

Self-supervised image enhancement research on low-light & medical imagery.

SSLImage EnhancementResearch
EPOCH_04loss ↓ 0.18
May 2024 — Aug 2024 · Open Source

Social Summer of Code · S3

Open Source Contributor

Shipped PRs across OSS CV / ML projects assigned through SSOC.

OSSPythonGit
EPOCH_03loss ↓ 0.31
Jun 2022 — Aug 2022 · Mumbai, India

Panel Technologies

Software Dev Intern

Built invoice digitization software — first taste of pipelines & document processing.

PythonAutomation
→ OUTPUT
CV ENGINEER · CONVERGED ✓
MOD_05// publications · neural co-authorship network

RESEARCH.GRAPH

graph.render · 4 papers · 1 author
NODE_V1In Proceeding
VigyaanSetu: AI-Powered Platform for Research Paper Understanding & Multimodal Knowledge Transformation
NODE_V2IEEE — In Proceeding
A Framework for AI-Driven Research Automation: The VigyaanSetu Approach
NODE_V3Springer Nature
Thermal & Mammographic Image Fusion for Breast Cancer Detection — Self-Supervised Bi-Pipeline
NODE_V4IEEE
Detecting Heart Disease Using Machine Learning
THESIS.STATEMENT  //  research_intent

I want to spend the next decade understanding how machines see, read and reason about the physical world — bridging low-level perception (depth, motion, segmentation) with high-level multimodal cognition (vision-language models, scene graphs, embodied agents). My long-term goal is a PhD focused on self-supervised & multimodal representation learning, with applications in medical imaging, assistive technology and scientific discovery.

ACTIVE.RESEARCH_INTERESTS  //  ranked
  • 01Self-supervised & contrastive representation learning
  • 02Multimodal LLMs · vision–language grounding · VQA
  • 033D perception · depth, NeRF, Gaussian splatting
  • 04Medical image fusion & diagnostic AI (mammo / thermal)
  • 05Assistive CV systems · wearable & edge inference
  • 06Research-automation agents · literature understanding
currently_reading: DINOv2 · LLaVA-NeXT · 4D-Gaussians · SAM 2