Skip to content
Work

Projects

Explore the open-source projects in full, then read about two private workplace agents and the problems they solved.

2025

ScrapeGPT

Self-hosted, AI-assisted web scraping

Paste a URL and an LLM proposes the extraction fields and selectors; the app re-validates and self-heals them against the real HTML, then crawls and exports clean CSV/JSON/XLSX.

FastAPIPostgreSQLLiteLLMReactPlaywright
2025

Aigram

Telegram, but AI-powered — a self-hosted messenger

Turns your own Telegram account into a self-hosted, installable AI messenger (formerly SakaiBot). Read and send messages with inline media in a glassy web app, and call an LLM right inside any chat — analyze, ask, translate, image, voice. AI results land in a panel, so you decide what to send.

PythonTelethonFastAPIPWAGemini
2026

PromptAmp

The prompt amplifier — your keys, any site

A browser extension that turns a rough draft into an engineered prompt inside any text field on any site — with your own API key. The result appears next to your original, nothing is replaced until you accept, and pasted code comes back byte-for-byte. Live on Firefox Add-ons.

TypeScriptWXTManifest V3VitestPlaywright
2025

RubricEval

A rubric-driven code-evaluation platform

Define a versioned rubric of weighted, gated criteria; submit a GitHub repo or ZIP and watch an LLM grade each criterion live against the real code — but a deterministic policy in code makes the final accept / review / reject call, reproducible and auditable.

Next.jsFastAPILiteLLMPostgreSQL
At work

Agents I shipped at work

Two private systems I designed and built end to end: a research agent that turns social feeds into source-linked briefs, and a multi-agent BI system that turns department data into checked, actionable reports.

Workplace agent

Social Research Agent

I designed and built a LangGraph research agent that turns the daily flood of posts into a short, source-linked brief — and answers focused questions on demand, like what a set of accounts or their followers are saying about a topic.

The problem

The team followed several fast-moving subjects on X (Twitter): topic categories, specific people, communities, and whatever was trending. Keeping up meant running the same searches by hand every day, wading through noise, and still missing the posts that mattered. Focused questions — what a group of accounts, or their followers, were saying about a topic — took even longer.

My role

I designed and built it end to end: the LangGraph workflow, the collectors for topic categories, people, followers, communities, and trends, the filtering and relevance logic, the LLM prompts, and a Docker deployment that produces the daily brief on a schedule.

How it works

  1. 01Scope
  2. 02Collect
  3. 03Filter & dedupe
  4. 04LLM relevance
  5. 05Summarize
  6. 06Deliver with source

Engineering highlights

Code decides the facts, the model decides relevance

Freshness, language, and duplicate checks are deterministic rules. The LLM is asked only what it's good at: is this worth a person's attention, and what topic is it?

Never the same finding twice

Links that were already delivered are skipped, and a thread is handled as one unit, so a post found through both a topic search and an account search appears in the brief once.

Briefs built to be verified

Every result carries a summary of up to 200 characters, two or three sentences on why it matters, and the original text with its link, so the reader can check the source before acting.

Testable, and resilient to flaky APIs

Data sources sit behind a small ports-and-adapters interface with a mock implementation, so the whole pipeline runs in tests without the network. Retry and cache layers keep a transient API failure from breaking the daily run.

Outcomes

  • A daily, source-linked brief replaced rounds of manual searching across categories, people, and communities.
  • Focused questions about a topic, a set of accounts, or their followers could be answered on demand instead of by hand.
  • Every finding was one click from its original, so the team could verify before acting.
Workplace agent

Business Intelligence Agents

I designed and built a multi-agent BI system: specialist agents for people operations, finance, and campaigns read department data through an MCP server, reason over figures that code has already checked, and write Persian reports with concrete next steps — then remember what the manager corrected.

The problem

Department data and past reports lived in different tools, and another set of fresh numbers wasn't what managers needed. They needed to know what had changed and why, what deserved attention, and what to do next — and for their corrections to carry into the next report instead of getting lost.

My role

I designed the architecture and built it end to end: an MCP server that exposes department data, report history, and saved feedback as tools; LangGraph orchestration across the specialist agents; a deterministic metrics layer; the Persian reporting prompts; and a scheduled Docker deployment that ran daily, weekly, and monthly analyses unattended.

How it works

  1. 01Fetch via MCP
  2. 02Validate & compute
  3. 03Specialist agents
  4. 04Persian report
  5. 05Manager review
  6. 06Feedback memory

Engineering highlights

Numbers first, narrative second

Code computes 20+ financial ratios and period-over-period changes before the model sees anything. The LLM interprets checked figures and never does the arithmetic.

What moves the business, not every line item

A materiality filter ignores bank fees and small charges, unusual transactions are flagged one by one, and costs are split into fixed and variable, so the finance report talks about what actually moves the business.

Each team measured by its own yardstick

The people-operations agent evaluates against the KPIs each team defined for itself; the campaign agent compares cost per view with its target and suggests where to move budget.

Memory you can inspect

Manager corrections are stored as explicit review notes and retrieved on the next run. No retraining, fully auditable, and the manager keeps the final say.

Outcomes

  • Managers received daily, weekly, and monthly reports per department — what changed, what needs attention, what to do next — without anyone assembling them by hand.
  • Recommendations rest on figures the code had already checked, not on the model's arithmetic.
  • Reports were published to shared spreadsheets in right-to-left Persian, ready for the team to read and discuss.
  • Each report built on the feedback given to the last one instead of starting from zero.