| AB-VietLLM | Language models | Compare models on Vietnamese tasks: quality, instruction following, tool use, speed and cost | Test cases, models, actual tokens |
| AB-Doc | Document processing | Accuracy of the important fields, effort to check and fix, processing time and cost | Pages processed, solutions compared, scoring configuration |
| AB-Voice | Voicebots | Task completion, serious errors, latency, interruptions, handover to a person | Simulated minutes and audio minutes evaluated, counted separately |
| AB-Memory | Long-term memory | Long context (measured on the base model) and recall across sessions (measured on the whole system) | Scenarios, sessions and turns, tokens |