We encourage you to evaluate how much Elicit helps with your use-case. You could evaluate:
How much time does Elicit save you in screening studies or extracting data?
How accurate is Elicit at extracting data from your studies?
What is Elicit's precision and recall for screening your studies?
Either use a random subset of data from an ongoing project, or use Elicit to replicate part of a systematic review that you've already finished.
Either way, we strongly recommend that when Elicit disagrees with human reviewers, you check who is right rather than assuming that the human review is right. In our testing, Elicit is often more accurate than humans.
Share your evaluation with us!
If you're working on an evaluation where you are comparing Elicit's performance to human research assistants or to another AI product and then publish the results, we'd love to hear about it. Please send evaluations to [email protected]!
