# Arsh Javed — Full Technical Context & Engineering Dossier > Comprehensive architecture documents, research paper methodologies, systems design decisions, and background for LLM and agentic crawlers. - Website: https://arshjaved.in/ - GitHub: https://github.com/ArshCypherZ - LinkedIn: https://www.linkedin.com/in/arshjaved/ - Hugging Face: https://huggingface.co/ArshCypherZ - Kaggle: https://www.kaggle.com/arshjaved - Telegram: https://t.me/halalneko --- ## 1. Biography & Professional Summary Arsh Javed is a Software Engineer and Computer Science undergraduate at Graphic Era University (Class of 2027) based in Dehradun, Uttarakhand, India. His work spans distributed systems engineering, high-concurrency event-driven backends, federated deep learning, and multi-modal machine learning. He is an accepted first-author researcher at peer-reviewed IEEE and international conferences, a hackathon champion (1st Place at IIT Guwahati Encode 2025; Top 5 Finalist out of 35,000+ at Google GenAI Exchange Hackathon), and the sole creator/maintainer of Emilia Bot, serving over 800,000 active users across 5,000+ communities. --- ## 2. Research & Academic Publications (First-Author) ### Paper 1: Reconstructing Global Models from Federated Client Weights - **Venue**: IC2E3 2026 · NIT Uttarakhand (International Conference on Computing, Communication and Energy Systems) - **Status**: Accepted, Presented, First-Author - **Problem Statement**: In non-IID federated learning environments, standard aggregation algorithms (such as McMahan's FedAvg) suffer from severe client drift and weight divergence when distributed client datasets exhibit extreme label skew and statistical heterogeneity. - **Methodology**: Researched and evaluated a learned aggregation mechanism combining four non-IID MentalBERT transformer classifiers. Client weight deltas are reconstructed and aggregated through a parameterized regression mapping rather than naive coordinate-wise averaging. - **Empirical Results**: - Test Set: 4,432 benchmark samples. - Achieved **93.50% test accuracy** and **R² = 0.953**, significantly outperforming standard Federated Averaging (FedAvg at **83.98%**) and approaching the upper-bound centralized training baseline (**94.04%**). - **Artifacts**: - Code: https://github.com/ArshCypherZ/MentalBert - Pretrained Model: https://huggingface.co/arshjaved/Final-FL - Training Notebooks: https://www.kaggle.com/code/arshjaved/novel - Certificate: https://files.arshjaved.in/u/bvRCSx.pdf ### Paper 2: Framing, Sentiment, and Agenda-Setting Bias in News Coverage - **Venue**: IEEE INDISCON 2026 · MNIT Jaipur (Flagship IEEE India Council International Conference) - **Status**: Accepted, Presented, First-Author - **Problem Statement**: Quantifying subtle ideological and political framing shifts across heterogeneous news outlets at scale requires multi-layered natural language processing combined with rigorous statistical validation rather than simple dictionary-based sentiment scoring. - **Methodology**: Analyzed a corpus of **58,289 articles** across 5 national news organizations over longitudinal publication cycles. Built an automated computational pipeline integrating RoBERTa (fine-grained sentiment polarity), BART (abstractive framing distillation), spaCy (entity co-occurrence networks), and BERTopic (unsupervised topic modeling). Applied non-parametric statistical hypothesis testing (Mann-Whitney U, Kruskal-Wallis) to validate systemic framing differences. - **Artifacts**: - Dataset & Code: https://www.kaggle.com/datasets/arshjaved/political-bias-analysis --- ## 3. Production Systems & Engineering Architecture ### 1. Leadly (http://leadly.live) - **Category**: Outbound Lead Intelligence & Social Selling Platform (SaaS) - **Core Stack**: Next.js, Express, BullMQ, Redis, Prisma, Docker, Dodo Payments, Gemini 1.5 - **Architecture**: - Continuously monitors high-intent Reddit subreddits and query clusters via Reddit's streaming API. - Asynchronous BullMQ worker queues decouple rate-limited upstream platform ingestion from token-metered downstream LLM generation. - Dual-pass Gemini classification pipeline: Pass 1 filters commercial buying intent from noise (scoring 0–100); Pass 2 drafts organic, subreddit-compliant replies avoiding spam detection. - Redis distributed locking prevents duplicate message processing across horizontally scaled worker instances. - Dodo Payments integration with automated tier provisioning and token quotas. ### 2. Emilia Bot (https://t.me/Elf_Robot) - **Category**: High-Concurrency Distributed Automation Platform for Telegram - **Scale**: 800,000+ active users, 5,000+ Telegram supergroups, 73+ loadable plugins - **Core Stack**: Python (asyncio), Pyrogram, Telethon, Redis, MongoDB, Groq AI LPUs - **Architecture**: - Employs a dual MTProto client driver (Pyrogram + Telethon) on a custom asyncio event loop to multiplex network sockets and handle burst message traffic exceeding 10,000 messages/min. - Multi-tier write-behind cache with Redis: frequent group permission checks, ban lists, and anti-flood configurations are served in under 1ms from RAM, reducing MongoDB IOPS by 95%. - Distributed cross-chat federation protocol: allows supergroup networks to synchronize ban lists instantly across thousands of chats during coordinated spam attacks. ### 3. Dialmate Live Voice Agent (https://github.com/arshcypherz/dialmate-backend) - **Award**: 1st Place Winner at IIT Guwahati Encode 2025 Hackathon - **Category**: Real-Time Full-Duplex Conversational Voice AI - **Core Stack**: FastAPI, Pipecat, Silero VAD, Twilio WebSockets, Deepgram, Cartesia TTS - **Architecture**: - Achieves sub-500ms voice round-trip latency (caller speech end to AI voice output). - Integrates Silero Voice Activity Detection (VAD) for natural interruption handling: callers can speak over the AI, instantly pausing audio synthesis and recalculating LLM context. - Twilio SIP trunking with bi-directional streaming WebSockets handles packet jitter and frame sequencing. - Live supervisor dashboard with real-time sentiment analytics, transcript streaming, and human operator intervention. ### 4. Exo-Checkmate (https://github.com/arshcypherz/Exo-Checkmate) - **Category**: Multi-Modal Deep Learning for Exoplanet Transit Verification - **Core Stack**: PyTorch, ONNX Runtime, scikit-learn, NASA TESS Photometric Data - **Architecture**: - Hybrid neural architecture combining: 1. 1D-CNN: extracts local transit ingress/egress morphology from folded photometric light curves. 2. Bi-LSTM with Temporal Attention: captures long-range orbital periodicity and highlights genuine planetary dips over stellar flares. 3. Physics-Informed MLP: injects stellar metadata (radius, surface gravity log(g), effective temperature). - Optimized with custom FocalLoss to overcome extreme 1:100+ positive-to-negative class imbalance. - ONNX Runtime export delivers 14x faster inference throughput for high-volume survey datasets. ### 5. ClarityHub (https://github.com/arshcypherz/ClarityHub) - **Category**: Document-to-Video Engine & Automated Pedagogical Studio - **Core Stack**: Next.js, FastAPI, Celery, Redis, Manim, Kokoro TTS, Gemini Flash - **Architecture**: - Transforms technical papers and markdown documentation into narrated 1080p video lessons. - Hierarchical semantic decomposition via Gemini parses complex docs into scene graphs and visual metaphors. - Sandboxed Python Manim execution engine with automated AST self-healing fixes vector animation syntax errors in real time. - Celery distributed worker tasks parallelize video frame rendering and audio alignment. ### 6. DrishtiMind (https://github.com/ArshCypherZ/DrishtiMind) - **Category**: AI Mental Wellness Toolkit Tailored for Indian Youth - **Core Stack**: Next.js 14, Tailwind CSS, Framer Motion, Prisma, PostgreSQL (Supabase RLS), Clerk - **Architecture**: - Interactive CBT (Cognitive Behavioral Therapy) reframing coach guiding users through cognitive distortions. - Longitudinal mood analytics, streak tracking, and encrypted private journaling with Row-Level Security. --- ## 4. Technical Arsenal & Proficiencies - **Languages**: Python, TypeScript, JavaScript, SQL, C++, HTML5, CSS3 - **Machine Learning & AI**: PyTorch, Hugging Face Transformers, ONNX Runtime, scikit-learn, Manim, LightGBM, RAG, Multi-Modal Systems, Federated Learning - **Backend & Distributed Systems**: FastAPI, Node.js, Express, Asyncio, Redis, BullMQ, Celery, WebSockets, WebRTC - **Databases**: PostgreSQL, MongoDB, Supabase, SQLite, Redis, Prisma ORM - **Cloud & DevOps**: Docker, Linux (Arch / Ubuntu), Git, Cloudflare Pages, Oracle Cloud Infrastructure (OCI Generative AI Certified) --- ## 5. Contact & Links - **Website**: https://arshjaved.in/ - **GitHub**: https://github.com/ArshCypherZ - **LinkedIn**: https://www.linkedin.com/in/arshjaved/ - **Telegram**: https://t.me/halalneko - **Hugging Face**: https://huggingface.co/ArshCypherZ - **Kaggle**: https://www.kaggle.com/arshjaved