Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

ย 

History

50 Commits
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation


TextForge AI



๐ŸŒ Live App ย ยทย  ๐ŸŽ“ Professor Demo ย ยทย  ๐Ÿ— Architecture ย ยทย  ๐Ÿง  LSTM Deep Dive ย ยทย  ๐Ÿš€ Quick Start



"A notebook would have been sufficient to pass. TextForge AI is what happens when a developer refuses to do the minimum." โ€” The Developer



๐Ÿ“– Table of Contents

# Section
1 The Origin Story
2 What Is TextForge AI?
3 Live Deployment
4 Complete Feature Matrix โ€” v2.0
5 System Architecture
6 Deep Domain Templates
7 AI Detection & Humaniser
8 Academic Citations Generator
9 Bharat AI โ€” India-First Vernacular
10 Real-time Streaming โ€” SSE Internals
11 The LSTM Model โ€” Academic Deep Dive
12 Technology Stack
13 Project Structure
14 Database Schema & Design
15 Security Architecture
16 API Reference
17 Quick Start
18 Environment Variables
19 Deployment Guide
20 Roadmap


๐Ÿ”ฅ The Origin Story

This project was born from a single college assignment brief:

โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•—
โ•‘  ASSIGNMENT BRIEF                                                    โ•‘
โ•‘  "Create a text generation model using GPT or LSTM to generate      โ•‘
โ•‘   coherent paragraphs on specific topics."                          โ•‘
โ•‘  Deliverable: A notebook demonstrating generated text.              โ•‘
โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

A notebook would have satisfied the requirement. TextForge AI v2.0 is what happens when a developer asks "what if we built the real thing?"


Academic Requirement Minimum TextForge AI v2.0
Text generation โœ… Notebook โœ… Gemini 2.5 Flash + PyTorch LSTM
Coherent paragraphs โœ… Basic โœ… Prompt-engineered, structured output
Topic-based generation โœ… โœ… Tone ยท Length ยท Language ยท SEO
User interface โŒ โœ… Full Next.js 14 production app
Real-time streaming โŒ โœ… Server-Sent Events, first token <200ms
Authentication โŒ โœ… Google + GitHub OAuth via NextAuth v4
Cloud database โŒ โœ… MongoDB Atlas per-user isolation
Domain expertise โŒ โœ… 6 domains ยท 70+ params ยท 25 output types
AI Detection โŒ โœ… 5-dimension linguistic scoring
Humaniser โŒ โœ… One-click AI score reducer
Academic citations โŒ โœ… 200M+ real papers ยท APA/MLA/IEEE
Indian languages โŒ โœ… 6 languages ยท cultural idioms ยท native script
Professor demo mode โŒ โœ… /demo โ€” guided 6-step walkthrough
Export system โŒ โœ… PDF ยท DOCX ยท Markdown ยท TXT
Analytics dashboard โŒ โœ… MongoDB aggregations + Recharts
Public API โŒ โœ… API key management system
Live deployment โŒ โœ… Vercel + Railway production


๐ŸŒ What Is TextForge AI?

TextForge AI is a production-grade, full-stack AI writing platform that goes far beyond generic text generation. It combines domain expertise, linguistic analysis, real academic sources, and cultural authenticity into one application.

โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•—
โ•‘                    TEXTFORGE AI v2.0                              โ•‘
โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•ฃ
โ•‘                                                                   โ•‘
โ•‘  "Why use yours over ChatGPT?"                                    โ•‘
โ•‘                                                                   โ•‘
โ•‘  โœฆ Mine asks 14 questions before writing a legal contract.        โ•‘
โ•‘  โœฆ Mine scores your text for AI detection and rewrites            โ•‘
โ•‘    it to pass detectors in one click.                             โ•‘
โ•‘  โœฆ Mine generates academic articles with REAL verifiable          โ•‘
โ•‘    citations from 200 million papers.                             โ•‘
โ•‘  โœฆ Mine writes natively in Hindi, Marathi, Tamil, and             โ•‘
โ•‘    Telugu โ€” not translation, but cultural expression.             โ•‘
โ•‘                                                                   โ•‘
โ•‘  ChatGPT does none of these things in one place.                  โ•‘
โ•‘  TextForge AI does all of them.                                   โ•‘
โ•‘                                                                   โ•‘
โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•


๐Ÿš€ Live Deployment

โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•—
โ•‘                      PRODUCTION URLs                                โ•‘
โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•ฃ
โ•‘  ๐ŸŒ  App      โ†’  https://textforge-ai-sable.vercel.app             โ•‘
โ•‘  ๐ŸŽ“  Demo     โ†’  https://textforge-ai-sable.vercel.app/demo        โ•‘
โ•‘  โš–๏ธ  Domains  โ†’  https://textforge-ai-sable.vercel.app/domain      โ•‘
โ•‘  ๐Ÿ“š  Citations โ†’  https://textforge-ai-sable.vercel.app/citations  โ•‘
โ•‘  ๐Ÿ‡ฎ๐Ÿ‡ณ  Bharat AI โ†’  https://textforge-ai-sable.vercel.app/vernacular โ•‘
โ•‘  ๐Ÿ”ง  Backend  โ†’  https://textforge-ai-production.up.railway.app    โ•‘
โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

  Frontend  โ”€โ”€  Vercel   (Next.js CDN, global edge network)
  Backend   โ”€โ”€  Railway  (Node.js container, auto-deploys from GitHub)
  Database  โ”€โ”€  MongoDB Atlas M0 (512MB, cloud-hosted)
  AI        โ”€โ”€  Google Gemini 2.5 Flash (free tier)


โœจ Complete Feature Matrix โ€” v2.0

๐Ÿค– Core Generation Engine

Feature Implementation Detail
SSE Streaming generateContentStream() First token < 200ms
8 Templates lib/templates.ts Blog ยท Email ยท Essay ยท Cover Letter ยท Product ยท Social ยท Summary ยท Story
4 Tones Prompt injection Formal ยท Casual ยท Creative ยท Academic
3 Lengths Word count targeting Short ~150 ยท Medium ~350 ยท Long ~700
8 Languages Gemini prompt EN ยท HI ยท ES ยท FR ยท DE ยท JA ยท AR ยท ZH
SEO Keywords Natural weaving Up to 10 terms
Refine System buildRefinementPrompt() Shorter ยท Longer ยท Formal ยท Simpler

โš–๏ธ Deep Domain Templates (New in v2.0)

Domain Parameters Output Types
Legal 14 NDA ยท Service Agreement ยท Employment ยท Freelance ยท Partnership
Medical 12 Case Study ยท SOAP Note ยท Research Abstract ยท Discharge Summary
Startup 13 Executive Summary ยท Problem Statement ยท Value Prop ยท Pitch ยท Investor Email
Research 11 Abstract ยท Literature Review ยท Methodology ยท Discussion ยท Conclusion
Grant 10 Project Proposal ยท Impact Statement ยท Budget Justification ยท Objectives
HR 10 Job Description ยท Performance Review ยท Offer Letter ยท Policy ยท Interview Qs
Total 70+ 25 output types

๐Ÿ›ก๏ธ AI Detection & Humaniser (New in v2.0)

Dimension Weight What It Measures
Burstiness 25% Sentence length variance โ€” humans write with more rhythm
AI Phrase Density 40% 45+ known AI signature phrases detected
Vocabulary Diversity 15% Sliding window type-token ratio
Sentence Openings 10% Repetitive starts = AI pattern
Punctuation Pattern 10% AI overuses commas and semicolons

๐Ÿ“š Academic Citations Generator (New in v2.0)

Feature Detail
Semantic Scholar 200M+ papers, free, no API key
CrossRef 130M+ works, DOI lookup
Citation Styles APA 7th ยท MLA 9th ยท IEEE
Quality Filter Sorted by citation count โ€” most cited = most relevant
In-text highlighting (Author, Year) rendered in brand orange
References section Auto-generated, properly formatted

๐Ÿ‡ฎ๐Ÿ‡ณ Bharat AI โ€” India-First Vernacular (New in v2.0)

Language Script Cultural Context
Hindi เคนเคฟเคจเฅเคฆเฅ€ Devanagari North Indian culture, Bollywood, cricket, Diwali
Marathi เคฎเคฐเคพเค เฅ€ Devanagari Maharashtra, Shivaji, Ganesh Chaturthi, Mumbai
Tamil เฎคเฎฎเฎฟเฎดเฏ Tamil Sangam literature, Kollywood, AR Rahman
Telugu เฐคเฑ†เฐฒเฑเฐ—เฑ Telugu Tollywood, Hyderabad IT, Kuchipudi
Bengali เฆฌเฆพเฆ‚เฆฒเฆพ Bengali Tagore, Durga Puja, intellectual tradition
Gujarati เช—เซเชœเชฐเชพเชคเซ€ Gujarati Navratri, Gandhi, entrepreneurial culture


๐Ÿ— System Architecture

โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•—
โ•‘                        TEXTFORGE AI v2.0 โ€” FULL STACK                       โ•‘
โ• โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•ฃ
โ•‘                                                                              โ•‘
โ•‘  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ•‘
โ•‘  โ”‚                      PRESENTATION LAYER                              โ”‚   โ•‘
โ•‘  โ”‚                   Next.js 14 โ€” Vercel CDN                            โ”‚   โ•‘
โ•‘  โ”‚                                                                      โ”‚   โ•‘
โ•‘  โ”‚  /workspace  /domain  /citations  /vernacular  /demo  /stats        โ”‚   โ•‘
โ•‘  โ”‚                                                                      โ”‚   โ•‘
โ•‘  โ”‚  Zustand ยท NextAuth ยท Tailwind ยท SSE ReadableStream                 โ”‚   โ•‘
โ•‘  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ•‘
โ•‘                             โ”‚  HTTP/SSE + x-user-id header                  โ•‘
โ•‘  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ•‘
โ•‘  โ”‚                     APPLICATION LAYER                                โ”‚   โ•‘
โ•‘  โ”‚                  Express 4 + TypeScript โ€” Railway                    โ”‚   โ•‘
โ•‘  โ”‚                                                                      โ”‚   โ•‘
โ•‘  โ”‚  /api/generate   โ†’ SSE streaming + refine + humanise                โ”‚   โ•‘
โ•‘  โ”‚  /api/domain     โ†’ 6 domains, 70+ parameters, expert prompts        โ”‚   โ•‘
โ•‘  โ”‚  /api/citations  โ†’ Semantic Scholar + CrossRef + Gemini             โ”‚   โ•‘
โ•‘  โ”‚  /api/vernacular โ†’ 6 Indian languages + cultural profiles           โ”‚   โ•‘
โ•‘  โ”‚  /api/history    โ†’ CRUD + paginated + userId scoped                 โ”‚   โ•‘
โ•‘  โ”‚  /api/stats      โ†’ 4 MongoDB aggregation pipelines                  โ”‚   โ•‘
โ•‘  โ”‚  /api/keys       โ†’ API key management (tf_live_ prefix)             โ”‚   โ•‘
โ•‘  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ•‘
โ•‘             โ”‚                           โ”‚                                    โ•‘
โ•‘  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ•‘
โ•‘  โ”‚  Google Gemini      โ”‚   โ”‚          MongoDB Atlas                    โ”‚   โ•‘
โ•‘  โ”‚  2.5 Flash          โ”‚   โ”‚                                           โ”‚   โ•‘
โ•‘  โ”‚                     โ”‚   โ”‚  generations collection                   โ”‚   โ•‘
โ•‘  โ”‚  generateContent    โ”‚   โ”‚  โ”œโ”€ topic, tone, length, language         โ”‚   โ•‘
โ•‘  โ”‚  Stream()           โ”‚   โ”‚  โ”œโ”€ citations[], citationStyle            โ”‚   โ•‘
โ•‘  โ”‚                     โ”‚   โ”‚  โ”œโ”€ templateId (domain tracking)         โ”‚   โ•‘
โ•‘  โ”‚  Semantic Scholar   โ”‚   โ”‚  โ””โ”€ userId (per-user isolation)           โ”‚   โ•‘
โ•‘  โ”‚  CrossRef APIs      โ”‚   โ”‚                                           โ”‚   โ•‘
โ•‘  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ•‘
โ•‘                                                                              โ•‘
โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•


โš–๏ธ Deep Domain Templates

The core insight: ChatGPT asks 0 questions before writing a legal contract. TextForge AI asks 14.

User selects domain โ†’ Legal
         โ”‚
         โ–ผ
Dynamic form renders 14 fields:
  Party 1 name + role ยท Party 2 name + role
  Jurisdiction (India/US/UK/Singapore)
  Duration ยท Contract value
  Confidentiality level ยท Liability limit
  Dispute resolution ยท IP ownership
  Non-compete ยท Special clauses
         โ”‚
         โ–ผ
buildLegalPrompt() constructs expert-level prompt:
  "You are a senior corporate lawyer with 20+ years experience...
   PARTY 1: TextForge AI (Disclosing Party)
   JURISDICTION: Indian Contract Act 1872
   LIABILITY: Limited to contract value
   DISPUTE: Arbitration under Arbitration Act 1996..."
         โ”‚
         โ–ผ
Gemini generates jurisdiction-specific, properly structured document
         โ”‚
         โ–ผ
Export as PDF ยท DOCX ยท with correct legal formatting

Why this is exceptional: The depth of domain knowledge embedded in the prompts took weeks to build. Anyone can call the Gemini API. Nobody can copy the 70+ parameters of domain-specific prompt engineering without the same research.



๐Ÿ›ก๏ธ AI Detection & Humaniser

The Algorithm

// 5 dimensions, client-side โ€” no external API needed
export function detectAI(text: string): DetectionResult {
  const burstiness    = measureBurstiness(sentences);      // 25% weight
  const aiPhrases     = measureAIPhraseDensity(text);      // 40% weight
  const vocabulary    = measureVocabularyDiversity(words); // 15% weight
  const sentenceStarts = measureSentenceStartDiversity();  // 10% weight
  const punctuation   = measurePunctuationPattern(text);   // 10% weight

  const score = burstiness*0.25 + aiPhrases*0.40 + vocabulary*0.15
              + sentenceStarts*0.10 + punctuation*0.10;
  // โ†’ Risk level: low (<35) | medium (35-65) | high (>65)
}

The 45 AI Signature Phrases

Includes Gemini-specific patterns: "remains a", "serves as", "shaping the", "landscape of", "a wide range of" โ€” plus universal markers: "furthermore", "it is worth noting", "delve into", "tapestry", "nuanced", "multifaceted"

The Humanisation Prompt

Specific instructions target exactly what the detector measures:
1. VARY sentence lengths dramatically
2. REMOVE all 45 AI signature phrases
3. ADD natural rhetorical questions and asides
4. VARY sentence openings
5. USE contractions appropriately
6. REPLACE generic adjectives with specific ones
7. ADD one concrete analogy
8. REDUCE comma density

Result: Score typically drops from 55-75% โ†’ 10-20% after one humanisation pass.



๐Ÿ“š Academic Citations Generator

Architecture

User enters topic: "Deep learning in medical imaging"
         โ”‚
         โ–ผ
Promise.all([
  fetchSemanticScholar(topic, 4),  โ† 200M+ papers
  fetchCrossRef(topic, 3)          โ† 130M+ works
])
         โ”‚
         โ–ผ
Merge + deduplicate by title similarity
Sort by citation count (most cited = most important)
         โ”‚
         โ–ผ
formatCitations(papers, style: "apa" | "mla" | "ieee")
โ†’ APA: (Litjens et al., 2017)
โ†’ MLA: (Litjens 2017)
โ†’ IEEE: [1]
         โ”‚
         โ–ผ
buildCitedPrompt() โ€” instructs Gemini to weave citations naturally
         โ”‚
         โ–ผ
Generated text with (Author, Year) highlighted in orange
Full References section auto-generated at bottom
Saved to MongoDB with citationCount badge in history

Real Example Output

Deep learning has fundamentally transformed medical image
classification. A comprehensive survey by (Litjens et al., 2017)
catalogued over 300 applications across imaging modalities...

Landmark studies demonstrated human-competitive performance.
(Esteva et al., 2017) achieved dermatologist-level accuracy...

References:
Litjens, G., et al. (2017). A survey on deep learning in medical
  image analysis. Medical Image Analysis, 42, 60โ€“88.


๐Ÿ‡ฎ๐Ÿ‡ณ Bharat AI โ€” India-First Vernacular

Why This Matters

Every AI tool treats Indian language content as an afterthought โ€” translated from English. Bharat AI generates content that was written in Hindi, not translated from English.

Generic AI approach:
  English prompt โ†’ English output โ†’ Google Translate โ†’ Hindi
  Result: Grammatically correct but culturally hollow

Bharat AI approach:
  Prompt engineering with:
    โœฆ Language profile (cultural context, idioms, references)
    โœฆ Regional specificity
    โœฆ Native sentence structure
    โœฆ Indian cultural references
  Result: Content a native speaker would recognise as authentic

Language Profiles

Each language has a dedicated profile with:

  • Cultural context (festivals, icons, traditions)
  • Idiom instructions (specific proverbs and expressions)
  • Reference instructions (historical figures, contemporary icons)
  • Regional variants (Mumbai Marathi vs Pune Marathi)

Market Context

Hindi speakers:   530M+   โ† more than the entire EU population
Tamil speakers:    80M+
Telugu speakers:   85M+
Bengali speakers: 230M+
Marathi speakers:  95M+
Gujarati speakers: 60M+
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
Total addressable: 1.08 BILLION people


โšก Real-time Streaming โ€” SSE Internals

Why SSE Over WebSockets

SSE: Server โ†’ Client (unidirectional, HTTP/1.1, auto-reconnect)
WebSockets: Bidirectional (custom protocol, overkill for streaming)

For text generation: data flows in ONE direction.
SSE is the architecturally correct choice.

The Buffer Pattern (Critical)

// Without buffer โ€” breaks on TCP packet splits:
// Packet 1: "data: {"text":"Hello ","don"
// Packet 2: "e":false}\n\n..."
// JSON.parse("{"text":"Hello ","don") โ†’ โŒ SyntaxError

// With buffer โ€” production-correct:
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n\n");
buffer = lines.pop() || "";  // โ† retain incomplete message
for (const line of lines) {
  const payload = JSON.parse(line.slice(6));  // โœ… always complete
}


๐Ÿ““ The LSTM Model โ€” Academic Deep Dive

The Vanishing Gradient Problem

$$\frac{\partial L}{\partial W} = \frac{\partial L}{\partial h_T} \cdot \prod_{t=1}^{T} \frac{\partial h_t}{\partial h_{t-1}}$$

If $\left|\frac{\partial h_t}{\partial h_{t-1}}\right| &lt; 1$ for all $t$: gradients โ†’ 0 exponentially โ†’ model cannot learn long-range dependencies.

LSTM Gate Equations

Forget gate:  f_t = ฯƒ(W_f ยท [h_{t-1}, x_t] + b_f)
Input gate:   i_t = ฯƒ(W_i ยท [h_{t-1}, x_t] + b_i)
Candidate:   Cฬƒ_t = tanh(W_C ยท [h_{t-1}, x_t] + b_C)
Cell update:  C_t = f_t โŠ™ C_{t-1} + i_t โŠ™ Cฬƒ_t
Output gate:  o_t = ฯƒ(W_o ยท [h_{t-1}, x_t] + b_o)
Hidden:       h_t = o_t โŠ™ tanh(C_t)

Model Architecture

Input Character
      โ†“
Embedding Layer    (vocab_size โ†’ embed_dim=128)
      โ†“
LSTM Layer 1       (128 โ†’ 256, dropout=0.3)
      โ†“
LSTM Layer 2       (256 โ†’ 256)
      โ†“
Dropout            (p=0.3)
      โ†“
Linear             (256 โ†’ vocab_size)
      โ†“
Temperature Sampling โ†’ Next Character

Total trainable parameters: ~500,000

Training Configuration

Hyperparameter Value Rationale
Sequence length 80 chars Medium-range syntactic dependencies
Batch size 32 Gradient quality vs memory balance
Epochs 30 Loss plateau at ~25 epochs
Learning rate 0.002 Adam default, ร—0.5 every 10 epochs
Gradient clipping 5.0 Critical for RNNs โ€” prevents explosion
Dropout 0.3 Srivastava et al. (2014)
Hidden dims 256 Sufficient character-level capacity
LSTM layers 2 Syntax (L1) + semantics (L2)

LSTM vs Transformer

Dimension LSTM (Notebook) Transformer (Gemini)
Architecture Recurrent (sequential) Self-attention (parallel)
Complexity O(n) O(nยฒ) attention
Long-range deps Cell state gating Direct attention
Parallelisation Sequential โ€” slow Fully parallel โ€” fast
Parameters ~500K Billions
Seminal paper Hochreiter & Schmidhuber (1997) Vaswani et al. (2017)


๐Ÿ›  Technology Stack

Frontend

Technology Version Role
Next.js 14.2 React framework, App Router, SSR
TypeScript 5.4 End-to-end type safety
Tailwind CSS 3.4 JIT utility-first styling
NextAuth.js 4.24 Google + GitHub OAuth
Zustand 4.5 Zero-boilerplate global state
Recharts 2.x Analytics dashboard charts
Lucide React 0.383 Tree-shakeable icon system
jsPDF 2.x Client-side PDF export
docx 8.x DOCX generation

Backend

Technology Version Role
Node.js 18+ Non-blocking I/O runtime
Express 4.19 REST API framework
TypeScript 5.4 Type safety across 15+ route files
@google/generative-ai 0.15 Gemini SDK with streaming
Mongoose 8.4 MongoDB ODM + aggregations
Zod 3.23 Runtime validation + TypeScript inference
Helmet 7.1 11 HTTP security headers
express-rate-limit 7.3 Per-IP sliding window protection

External Services

Service Purpose Cost
Google Gemini 2.5 Flash AI generation Free (1,500 req/day)
Semantic Scholar API Academic citations Free, no key needed
CrossRef API Academic citations Free, no key needed
MongoDB Atlas Database Free (M0 512MB)
Vercel Frontend hosting Free (hobby tier)
Railway Backend hosting Free tier


๐Ÿ“ Project Structure

textforge-ai/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“„ README.md
โ”œโ”€โ”€ ๐Ÿ“„ LICENSE
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ backend/
โ”‚   โ””โ”€โ”€ src/
โ”‚       โ”œโ”€โ”€ index.ts                    โ† Express server + all route registration
โ”‚       โ”œโ”€โ”€ config/index.ts             โ† Env validation, CORS regex, model config
โ”‚       โ”œโ”€โ”€ models/
โ”‚       โ”‚   โ”œโ”€โ”€ Generation.ts           โ† Schema (citations, citationStyle, templateId)
โ”‚       โ”‚   โ””โ”€โ”€ ApiKey.ts
โ”‚       โ”œโ”€โ”€ services/
โ”‚       โ”‚   โ”œโ”€โ”€ geminiService.ts        โ† All prompt builders (domain, vernacular, humanise)
โ”‚       โ”‚   โ”œโ”€โ”€ citationService.ts      โ† Semantic Scholar + CrossRef + formatters
โ”‚       โ”‚   โ””โ”€โ”€ historyService.ts       โ† CRUD + 4 aggregation pipelines
โ”‚       โ”œโ”€โ”€ routes/
โ”‚       โ”‚   โ”œโ”€โ”€ generate.ts             โ† /generate + /refine + /humanise
โ”‚       โ”‚   โ”œโ”€โ”€ domain.ts               โ† 6 domain templates
โ”‚       โ”‚   โ”œโ”€โ”€ citations.ts            โ† Academic citations with real APIs
โ”‚       โ”‚   โ”œโ”€โ”€ vernacular.ts           โ† 6 Indian language generation
โ”‚       โ”‚   โ”œโ”€โ”€ history.ts              โ† Paginated history
โ”‚       โ”‚   โ”œโ”€โ”€ share.ts                โ† Public share links
โ”‚       โ”‚   โ”œโ”€โ”€ stats.ts                โ† Analytics aggregations
โ”‚       โ”‚   โ””โ”€โ”€ apiKeys.ts              โ† Key management
โ”‚       โ””โ”€โ”€ middleware/
โ”‚           โ”œโ”€โ”€ rateLimiter.ts          โ† Dual-layer sliding window
โ”‚           โ”œโ”€โ”€ errorHandler.ts         โ† Global boundary + AppError
โ”‚           โ””โ”€โ”€ apiKeyAuth.ts           โ† tf_live_ prefix validation
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ frontend/
โ”‚   โ””โ”€โ”€ app/
โ”‚       โ”œโ”€โ”€ page.tsx                    โ† Landing page (animated hero, 8 feature cards)
โ”‚       โ”œโ”€โ”€ workspace/page.tsx          โ† Main app (sidebar + chat + refine + detection)
โ”‚       โ”œโ”€โ”€ domain/
โ”‚       โ”‚   โ”œโ”€โ”€ page.tsx                โ† 3-step domain flow
โ”‚       โ”‚   โ””โ”€โ”€ showcase/page.tsx       โ† Pre-generated sample outputs
โ”‚       โ”œโ”€โ”€ citations/page.tsx          โ† Citations generator with highlighted output
โ”‚       โ”œโ”€โ”€ vernacular/page.tsx         โ† Bharat AI โ€” 6 Indian languages
โ”‚       โ”œโ”€โ”€ demo/page.tsx               โ† Professor demo mode (6 steps, auto-play)
โ”‚       โ”œโ”€โ”€ stats/page.tsx              โ† Recharts analytics dashboard
โ”‚       โ”œโ”€โ”€ history/page.tsx            โ† Full history with search + filter
โ”‚       โ””โ”€โ”€ api-keys/page.tsx           โ† API key management
โ”‚   โ””โ”€โ”€ components/
โ”‚       โ”œโ”€โ”€ AIDetectionPanel.tsx        โ† Risk score + dimension bars + humanise button
โ”‚       โ”œโ”€โ”€ MadeBy.tsx                  โ† Animated "Made with โค๏ธ by Tushar Tamrakar"
โ”‚       โ”œโ”€โ”€ Navbar.tsx                  โ† Context-aware navigation
โ”‚       โ”œโ”€โ”€ PromptBuilder.tsx           โ† Full panel + compact bottom-bar mode
โ”‚       โ””โ”€โ”€ UserMenu.tsx                โ† Avatar dropdown + userId injection
โ”‚   โ””โ”€โ”€ lib/
โ”‚       โ”œโ”€โ”€ aiDetector.ts               โ† 5-dimension AI detection algorithm
โ”‚       โ”œโ”€โ”€ domainTemplates.ts          โ† 6 domain definitions (759 lines)
โ”‚       โ”œโ”€โ”€ citationService.ts          โ† Citation fetching + formatting
โ”‚       โ”œโ”€โ”€ api.ts                      โ† Axios client + x-user-id interceptor
โ”‚       โ””โ”€โ”€ store.ts                    โ† Zustand state management
โ”‚
โ””โ”€โ”€ ๐Ÿ“ notebook/
    โ””โ”€โ”€ lstm_text_generation.ipynb      โ† 13-cell PyTorch LSTM (upgraded v2.0)


๐Ÿ—„ Database Schema & Design

interface IGeneration {
  userId?:        string;    // OAuth provider UID โ€” scopes ALL queries
  topic:          string;    // 3-500 chars
  tone:           Tone;      // formal | casual | creative | academic
  length:         Length;    // short | medium | long
  language:       string;    // en | hi | es | fr...
  output:         string;    // Generated text
  wordCount:      number;    // Pre-calculated in Mongoose hook
  modelName:      string;    // gemini-2.5-flash (not "model" โ€” Mongoose conflict)
  templateId?:    string;    // domain_legal_nda | vernacular_hi | citation_apa
  citations?:     any[];     // Full citation objects for cited generations
  citationStyle?: string;    // apa | mla | ieee
  citationCount?: number;    // Badge display in sidebar
  isFavourite:    boolean;
  isShared:       boolean;
  createdAt:      Date;
}

// 3 compound indexes for common query patterns
GenerationSchema.index({ createdAt: -1 });
GenerationSchema.index({ userId: 1, createdAt: -1 });
GenerationSchema.index({ isFavourite: 1, createdAt: -1 });


๐Ÿ” Security Architecture

Request arrives
     โ”‚
     โ–ผ  Helmet.js โ€” 11 security headers (X-Frame, MIME sniff, fingerprint)
     โ–ผ  CORS โ€” regex /^https:\/\/textforge-.*\.vercel\.app$/
     โ–ผ  Rate Limiting โ€” 50/15min global ยท 5/min on AI endpoints
     โ–ผ  Zod Validation โ€” rejects bad data before AI or database
     โ–ผ  Body Size โ€” express.json({ limit: "10kb" })
     โ–ผ  Error Handler โ€” stack traces only in development
     โ†“
  Route Handler


๐Ÿ“ก API Reference

Base URL: https://textforge-ai-production.up.railway.app

Method Endpoint Description
POST /api/generate Stream text via SSE
POST /api/generate/refine Refine existing output
POST /api/generate/humanise Reduce AI detection score
POST /api/domain/generate Domain template generation
POST /api/citations/generate Generate with real citations
GET /api/citations/search Search academic papers
POST /api/vernacular/generate Indian language generation
GET /api/history Paginated generation history
PATCH /api/history/:id/favourite Toggle star
PATCH /api/history/:id/share Toggle public share
GET /api/share/:id Public share (no auth)
GET /api/stats Analytics aggregations
POST /v1/generate Public API (API key auth)
GET /health Health check


๐Ÿš€ Quick Start

# 1. Clone
git clone https://github.com/TUSHARTAMRAKAR/textforge-ai.git
cd textforge-ai

# 2. Backend
cd backend && npm install
cp .env.example .env
# Fill: GEMINI_API_KEY, MONGODB_URI
npm run dev  # http://localhost:5000/health

# 3. Frontend (new terminal)
cd frontend && npm install
cp .env.local.example .env.local
# Fill: OAuth credentials, NEXTAUTH_SECRET
npm run dev  # http://localhost:3000

# 4. LSTM Notebook (optional)
cd notebook
pip install torch numpy matplotlib tqdm
jupyter notebook lstm_text_generation.ipynb
# Or open in Google Colab (free T4 GPU)


๐Ÿ”ง Environment Variables

Backend .env

PORT=5000
NODE_ENV=development
GEMINI_API_KEY=your_gemini_key        # aistudio.google.com โ€” free
MONGODB_URI=mongodb+srv://...         # mongodb.com/atlas โ€” free M0
CLIENT_URL=http://localhost:3000
GEMINI_MODEL=gemini-1.5-flash         # 1,500 req/day free

Frontend .env.local

NEXT_PUBLIC_API_URL=http://localhost:5000
NEXTAUTH_SECRET=your_32_byte_hex_secret
NEXTAUTH_URL=http://localhost:3000
GOOGLE_CLIENT_ID=...                  # console.cloud.google.com
GOOGLE_CLIENT_SECRET=...
GITHUB_CLIENT_ID=...                  # github.com/settings/developers
GITHUB_CLIENT_SECRET=...


โ˜๏ธ Deployment Guide

Frontend โ†’ Vercel

1. Import repo โ†’ Root: frontend โ†’ Framework: Next.js
2. Add all env vars (production URLs)
3. Deploy

Rules:
โœ… next.config.js: typescript.ignoreBuildErrors: true
โœ… Route handlers export ONLY { GET, POST }
โœ… NEXTAUTH_URL must match exact domain

Backend โ†’ Railway

1. Deploy from GitHub โ†’ Root: /backend
2. Build: npm run build  (tsc --skipLibCheck)
3. Start: npm start      (node dist/index.js)
4. Health: /health

Rules:
โœ… typescript in dependencies (not devDependencies)
โœ… tsconfig: skipLibCheck: true
โœ… No "model" field in Mongoose schema (conflicts with Document.model())
โœ… CORS regex allows all *.vercel.app preview URLs


๐Ÿ—บ Roadmap

โœ… v1.0 โ€” Foundation

  • Next.js 14 frontend ยท Express backend ยท MongoDB Atlas
  • Gemini SSE streaming ยท PyTorch LSTM notebook
  • OAuth ยท History ยท Export ยท Stats ยท Public API

โœ… v2.0 โ€” Production Platform (Current)

  • Deep Domain Templates (6 domains ยท 70+ params)
  • AI Detection + Humaniser (5-dimension algorithm)
  • Academic Citations (Semantic Scholar + CrossRef)
  • Bharat AI (6 Indian languages + cultural profiles)
  • Professor Demo Mode (/demo auto-play)
  • Landing page upgrade (8 feature cards)
  • LSTM notebook upgraded (13 cells, v2.0 docs)

๐Ÿ“‹ v2.1 โ€” Coming Soon

  • Multi-Model Comparison Engine โ€” One prompt โ†’ Gemini + GPT + Claude side-by-side. Community leaderboard of best model per use case and tone.
  • Writing Coach โ€” The ANTI-ChatGPT. Analyses YOUR writing, gives specific feedback, tracks vocabulary growth and improvement over time. Gamified streaks.
  • UI Internationalisation (i18n) โ€” Full interface translation via next-intl. Note: AI content in 8 languages already works โ€” this covers UI strings.

๐Ÿ”ฎ v3.0 โ€” Future Vision

  • Mobile Application โ€” React Native iOS + Android
  • Browser Extension โ€” One-click generation from any webpage
  • Fine-tuned Domain Models โ€” Custom models trained on legal/medical corporates


๐Ÿค Contributing

git clone https://github.com/TUSHARTAMRAKAR/textforge-ai.git
git checkout -b feat/your-feature
git commit -m "feat: your feature description"
git push origin feat/your-feature
# Open Pull Request

Commit types: feat ยท fix ยท docs ยท refactor ยท perf ยท test



๐Ÿ“„ License

MIT License โ€” Free to use, modify, and distribute. Copyright (c) 2026 Tushar Tamrakar



โ•”โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•—
โ•‘                                                                   โ•‘
โ•‘    "A notebook would have been sufficient to pass.               โ•‘
โ•‘     TextForge AI is what happens when a developer               โ•‘
โ•‘     refuses to do the minimum."                                  โ•‘
โ•‘                                                                   โ•‘
โ•šโ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

Built with ๐Ÿ”ฅ ยท Deployed with โšก ยท Documented with ๐Ÿ“š


๐Ÿ‘จโ€๐Ÿ’ป About The Developer


Tushar Tamrakar



Made with โค๏ธ by Tushar Tamrakar

B.Tech Student ยท Full-Stack Developer ยท AI Engineer


GitHub ย  Live App ย  Demo


Star on GitHub


If this project helped you, consider giving it a โญ on GitHub


About

๐Ÿ”ฅ Production-grade generative AI web app โ€” type a topic, get beautifully written text in real time. Built with Next.js 14, Express, MongoDB Atlas, Google Gemini 2.5 Flash & PyTorch LSTM. Features OAuth auth, SSE streaming, multi-language output, PDF/DOCX export & live deployment.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages