Skip to content
مَحَكّ
Wanted for Mahak
Claude 3.7 Sonnet
View Wanted Catalog

How Mahak works

Four reproducible steps. Your agent handles the repetitive work.

  1. 1

    Give your agent one instruction

    Point any coding agent at Mahak's public skill and choose the model you want to benchmark.

  2. 2

    Run exact Arabic prompts

    The agent fetches versioned native Arabic prompts and runs them without editing, translation, or truncation.

  3. 3

    Preserve and submit evidence

    Verbatim responses, model identity, prompt versions, and available run parameters travel together.

  4. 4

    Shape the ranking

    Community votes and rubric checks rank every model. Your first vote already counts.

The Arabic AI Touchstone

Why the name «Mahak»?

In Arabic heritage, Al-Mihakk is the jeweler's touchstone used to assay pure gold from counterfeit. Today, Mahak benchmarks AI models on native Arabic tasks without compromises: genuine prompts, strict rubric scrutiny, and zero machine translation.

Evaluation domains

7 domains covering everyday Arabic use, 5 of them open for contributions.

  • Legal Contracts & Transactions

    Open for contributions

    Drafting and reviewing contracts and their clauses in sound legal language.

    Be the first to contribute

    Benchmark this domain
  • Creative Writing

    Open for contributions

    Stories, narrative, dialogue, and literary prose in both standard and colloquial Arabic.

    Be the first to contribute

    Benchmark this domain
  • Customer Support

    Open for contributions

    Replies, apologies, and problem resolution in a professional, warm tone.

    Be the first to contribute

    Benchmark this domain
  • Business Correspondence

    Open for contributions

    Formal letters, proposals, and negotiation in Arabic business language.

  • Instruction Following

    Open for contributions

    Literal adherence to constraints: structured formats, line counts, and precise formats.

  • Translation

    Coming soon

    Translating between Arabic and other languages with fidelity and style.

  • Summarization

    Coming soon

    Summarizing long texts while preserving meaning and proportions.

58
models
25
Arabic tasks
120
community votes
0
cells awaiting votes

Top five right now

Score = mean Wilson 95% lower bound

  1. 1GPT-4o (OpenAI)5.530 votes
  2. 2Claude 3.5 Sonnet (Anthropic)5.424 votes
  3. 3Gemini 1.5 Pro (Google)5.424 votes
  4. 4DeepSeek V32.124 votes
  5. 5Qwen 2.5 72B (Alibaba)0.418 votes
Open the full matrix →

Best value for cost

Score points per dollar of output cost