Mahmoud Zaher

I'm Mahmoud Zaher, a machine learning engineer based in Cairo. I build systems that read messy, real-world documents — and I care about making them reliable, trustworthy, and explainable, not just accurate. My work spans document AI, model interpretability, and vision-language models, with earlier roots in computer vision, NLP, and speech synthesis.

Now

AI Engineer at Click ITS (Cairo, Dec 2025 – present), where I work on the document side of an enterprise agentic platform — making the system read difficult Arabic documents reliably, since that's the foundation everything else depends on.

  • Fine-tuned vision-language models for Arabic financial document extraction — reducing CER from >18% to <8% and lifting field-level accuracy to >0.92.
  • Built a three-stage synthetic data generation pipeline for OCR and layout-analysis VLM training, because good training data didn't exist.
  • Improved RAG recall via document expansion with query prediction, grounded in a domain-specific glossary.
  • Built and deployed an internal annotation platform used across teams.

Research

Interpretable AI — Remote intern supervised by a senior PhD researcher at the University of Queensland, Australia (Aug 2025 – Jan 2026).

  • Applied gradient-based feature attribution to pre-trained vision and language models.
  • Studied how saliency maps and weight structure evolve under post-hoc vs. intrinsic interpretability paradigms.
  • Investigated the grokking phenomenon — how networks acquire predictive competence and generalization.

Expressive Arabic TTS — Remote intern, ReachSci STEM Mini-PhD Programme (Oct 2023 – Apr 2024). Led a literature review of state-of-the-art generative TTS and fine-tuned GlowTTS and VITS to improve Arabic synthesis quality.

Selected experience

  • Full-Stack Developer, ITI Intensive Code Camp (Jul – Dec 2025) — built and shipped full-stack apps; automated build, test, and deployment with Jenkins, Docker, Linux, and CI/CD.
  • Digital-twin modeling for naval propulsion (R&D freelance)(Jul – Dec 2024) — engineered cost/degradation features from CODLAG telemetry, handled multicollinearity (VIF, PCA, regularization), and benchmarked maintenance-cost regressors.
  • Computer Vision & Deep Learning Intern, ITI (2022) — image classification, object detection, and semantic segmentation.

Projects

  • ActveX — co-founder & ML engineer. Fine-tuned YOLOv8m-Pose for real-time posture detection/correction on a self-collected dataset. Deployed, used by real users, and selected for a government incubation program.
  • Readify (graduation project) — team lead & ML engineer. Built an Arabic PDF/EPUB extraction pipeline (CER <15%), curated a ~25-hour Arabic speech corpus, and fine-tuned XTTS-v2 for speaker adaptation.

Education & recognition

BSc, Computer & Control Systems Engineering, University of Mansoura (2019 – 2024) — GPA 3.71 / 4.0.

  • Ministry of Youth & Sports incubation program (with ASRT), for ActveX.
  • 3rd place, 2024 Arab IoT & AI Challenge.
  • Qualified to present at GITEX Global 2024, Dubai.

Toolkit

Python, PyTorch, Transformers, LangChain, FastAPI/Django, React, TypeScript, LaTeX; PostgreSQL, MongoDB, ChromaDB, Neo4j; Docker, Jenkins, Linux, Git, CI/CD. Native Arabic, fluent English, basic French.

Get in touch

Reach me via the contact form, by email at maredazaher@gmail.com, or on GitHub, X, LinkedIn. I'm actively looking for challenging problems in trustworthy and explainable ML, and open to graduate research opportunities.