Joshua Smith Joshua Smith

Why Models Hallucinate

This might also have implications outside of new models. Including training data multiple times for an output to increase the likelihood it catches the correct answer might be a decent strategy. Hopefully the files won't get too big and cause other issues, but for generating a timeline for instance, having multiple copies of medical charts might increase accuracy without additional time or resources.

Assuming that the increased size fits within the token limit so that you do not run into other issues... (here's hoping token limits will be a think of the past soon! Maybe the trillions in infrastructure will help:/)

Source: https://cdn.openai.com/pdf/d04913be-3f6f-4d2b-b283-ff432ef4aaa5/why-language-models-hallucinate.pdf

Read More
Joshua Smith Joshua Smith

Agents vs Models

I have found that using advanced agentic like prompts directly into GPT5 over 4o agent mode in Copilot shows a dramatic difference in the quality of the output. In my workflow, it was able to generate numerous litigation discovery documents with extremely few errors that were mostly stylistic or formatting. GPT5 can reference hundreds of pages of documents across multiple providers, and formats in order to complete these tasks.

The experience of using it has almost been a bit spooky. The depth of knowledge that it has on tap is extroadinary, and now the hallucinations are down to a useable level in certain usecases with proper review techniques. I've been developing the new company website and wanted to have a knowledge base for people to go to thats focused on business owners with an action link at the bottom if a reader has that sort of case. But, it was spooky because it must have lifted some of its tips from m&a textbooks and biographies. GPT5 was using ebitda optimazation strategies that I learned working at Peregrine... The research and strategy implications over prior models can not be understated.

It can also read medical charts with extreme accuracy, with some results as low as a 1.6% error rate... that is almost expert level review in seconds which can be scheduled the moment you get access to the records.

Source: https://mpgone.com/is-gpt-5-accurate/

Read More
Joshua Smith Joshua Smith

How long does it take?

What an interesting take, AI has long been limited by dataset and compute resources, but now this new framework will help to implement more practical uses for AI.

I'm excited for the day when every employee is a manager of several AI agents. While some expect AI to enable lazy employees, it can also increase productivity to new heights.

Source: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/

Read More