Mahmoud Zaher
I'm Mahmoud Zaher, a machine learning engineer based in Cairo. I build systems that read messy, real-world documents — and I care about making them reliable, trustworthy, and explainable, not just accurate. My work spans document AI, model interpretability, and vision-language models, with earlier roots in computer vision, NLP, and speech synthesis.
Now
AI Engineer at Click ITS (Cairo, Dec 2025 – present), where I work on the document side of an enterprise agentic platform — making the system read difficult Arabic documents reliably, since that's the foundation everything else depends on.
- Fine-tuned vision-language models for Arabic financial document extraction — reducing CER from >18% to <8% and lifting field-level accuracy to >0.92.
- Built a three-stage synthetic data generation pipeline for OCR and layout-analysis VLM training, because good training data didn't exist.
- Improved RAG recall via document expansion with query prediction, grounded in a domain-specific glossary.
- Built and deployed an internal annotation platform used across teams.
Research
Interpretable AI — Remote intern supervised by a senior PhD researcher at the University of Queensland, Australia (Aug 2025 – Jan 2026).
- Applied gradient-based feature attribution to pre-trained vision and language models.
- Studied how saliency maps and weight structure evolve under post-hoc vs. intrinsic interpretability paradigms.
- Investigated the grokking phenomenon — how networks acquire predictive competence and generalization.
Expressive Arabic TTS — Remote intern, ReachSci STEM Mini-PhD Programme (Oct 2023 – Apr 2024). Led a literature review of state-of-the-art generative TTS and fine-tuned GlowTTS and VITS to improve Arabic synthesis quality.
Selected experience
- Full-Stack Developer, ITI Intensive Code Camp (Jul – Dec 2025) — built and shipped full-stack apps; automated build, test, and deployment with Jenkins, Docker, Linux, and CI/CD.
- Digital-twin modeling for naval propulsion (R&D freelance)(Jul – Dec 2024) — engineered cost/degradation features from CODLAG telemetry, handled multicollinearity (VIF, PCA, regularization), and benchmarked maintenance-cost regressors.
- Computer Vision & Deep Learning Intern, ITI (2022) — image classification, object detection, and semantic segmentation.
Projects
- ActveX — co-founder & ML engineer. Fine-tuned YOLOv8m-Pose for real-time posture detection/correction on a self-collected dataset. Deployed, used by real users, and selected for a government incubation program.
- Readify (graduation project) — team lead & ML engineer. Built an Arabic PDF/EPUB extraction pipeline (CER <15%), curated a ~25-hour Arabic speech corpus, and fine-tuned XTTS-v2 for speaker adaptation.
Education & recognition
BSc, Computer & Control Systems Engineering, University of Mansoura (2019 – 2024) — GPA 3.71 / 4.0.
- Ministry of Youth & Sports incubation program (with ASRT), for ActveX.
- 3rd place, 2024 Arab IoT & AI Challenge.
- Qualified to present at GITEX Global 2024, Dubai.
Toolkit
Python, PyTorch, Transformers, LangChain, FastAPI/Django, React, TypeScript, LaTeX; PostgreSQL, MongoDB, ChromaDB, Neo4j; Docker, Jenkins, Linux, Git, CI/CD. Native Arabic, fluent English, basic French.
Get in touch
Reach me via the contact form, by email at maredazaher@gmail.com, or on GitHub, X, LinkedIn. I'm actively looking for challenging problems in trustworthy and explainable ML, and open to graduate research opportunities.