Autonomous Telegram AI Assistant

"Brian AI" Telegram Bot

An autonomous multimodal Telegram personal assistant powered by n8n, Google Gemini 1.5, and session memory buffers for real-time voice and text intelligence.

System Architecture & Workflow Canvases

Brian Ai Telegram Personal Assistant

Brian Ai Telegram Personal Assistant
🔍 Zoom View

End-to-end n8n production workflow integrating Telegram webhooks, MIME audio detection, Google Gemini speech-to-text, and session-buffered AI reasoning.

Case Study Breakdown

Introduction

Brian AI — Autonomous Telegram Multi-Modal Personal Assistant

Mobile team members frequently alternate between sending quick text queries and recording voice memos while in transit, creating communication bottlenecks, slow manual playback overhead, and fragmented conversation context.

The Challenges

  • On-The-Go Voice Memo Bottleneck: Listening to incoming audio memos, transcribing action items, and typing manual replies caused severe delays of 15 to 45 minutes per voice note.
  • Stateless Bots & Context Loss: Traditional Telegram bots treat every message as an isolated webhook, forcing users to repeatedly re-explain background details with every follow-up question.
  • Cost & Latency Inefficiencies: Naive bot setups either ran every message through expensive transcription APIs or relied on fragmented separate bots for voice and text, multiplying costs and response lag.

The Solution

Designed and deployed an intelligent, end-to-end multimodal automation workflow inside n8n, combining Google Gemini 1.5, smart MIME routing, and session memory buffers:

  • Intelligent Ingestion & Selective Routing: Listens live via Telegram webhooks and checks MIME types instantly, bypassing transcription for text messages to save costs while routing voice notes to the audio pipeline.
  • Gemini Audio STT & Stream Aggregation: Downloads raw Telegram voice notes and transcribes them in seconds via Google Gemini STT, merging transcripts and standard text into a unified input variable.
  • Context-Aware AI Agent & Session Memory: Connects the Google Gemini 1.5 Chat Model to a window buffer Simple Memory node, tracking isolated user session IDs for natural multi-turn conversations delivered directly into Telegram.

The Results

  • Sub-30s Turnaround Speed: Cut message handling time from 15–45 minutes down to under 30 seconds for both voice memos and written queries.
  • 50%+ API Cost Savings: Selective payload routing eliminated wasted speech-to-text calls on standard text traffic, substantially reducing token overhead.
  • 24/7 Contextual Continuity: Established an always-on personal assistant with zero memory collisions, giving users instant recall across multi-step conversations anytime.
Project Information
Title Workflow: "Brian AI" Telegram Bot
Executed: May 2026
IMPACT

Modernized mobile Telegram communications with an autonomous multimodal AI pipeline, slashing response time to under 30s while cutting API transcription costs by over 50%.

Automation Specialist: Kent Brian Jintapa
In Collaboration: Errol John Peusca, Mark Dalumpines, and Joyce Ann Yap