-
You can't judge an AI agent until you know what the task is worth
We keep asking whether AI coding agents are good or bad, fast or slow, cheap or expensive. But good compared to what? Until you know what a task is actually worth to you, you have no way to tell whether the agent's output was a bargain or a rip-off.
-
ProxyStat is now available for Windows
Use ProxyStat to easily see if you have system proxy configured on Windows
-
ProxyStat v1.2.1 - Quick access to system proxy settings
ProxyStat v1.2.1 adds quick access to system proxy settings and is now notarized for easier setup on macOS
-
Benchmark models using OpenAI-compatible APIs
Learn to benchmark language models with OpenAI-compatible APIs using our updated Jupyter Notebook for optimal performance evaluation.
-
Language model benchmarks only tell half a story
When it comes to language models, we tend to look at benchmarks to decide which model is the best to use in our application. But benchmarks only tell half a story. Unless you're building an all-purpose chat application, what you should be actually looking at is how well a model works for your application.