Original Reddit post

I am testing whether AI can do useful work that requires reasoning across a large and technically complex set of documents. For this test, I used a real patent dispute called IPR2025-00030 . Patent cases are useful for this purpose because the record can include legal arguments, technical documents, expert testimony, and earlier patents. The information needed to reach a conclusion is spread across many pages, and the patent judges eventually publish a detailed decision that can be used as a reference. The case involved a patent related to power management in radio-frequency systems. One side argued that all 20 claims in the patent should be found unpatentable. I gave the AI the public case record that existed before the final ruling and asked it to write its own complete decision. The AI did not have access to the official decision, its later correction, documents added after the cutoff date, the internet, or outside information. I saved the AI’s answer before showing it the official result. The AI and the patent judges reached opposite conclusions: The AI concluded that none of the 20 claims had been shown to be unpatentable. The corrected official decision concluded that all 20 claims were unpatentable. I then showed the official decision to the same AI and asked it to compare the two decisions. The AI acknowledged that it matched the official result on zero of the 20 claims and zero of the six main arguments. Despite that, it concluded that its own reasoning was stronger overall. The central disagreement involved power efficiency. In simple terms, the official decision accepted measurements involving voltage, current, and radio output as evidence supporting the required power-efficiency behavior. The AI argued that this evidence did not clearly prove the required relationship between the power entering the system and the useful power leaving it. This leaves two questions: Was the AI’s original reasoning sound, despite reaching the opposite result? Did the AI compare the two decisions fairly, or did it defend the same mistake it had already made? The second question matters if we want to use AI to evaluate AI-generated work. Expert review is expensive, so using AI as a judge could make larger benchmarks possible. But that approach may not work if the AI prefers its own earlier reasoning. I am not trained in patent law, so I cannot reliably answer these questions myself. I am sharing the complete materials so people with relevant legal or technical knowledge can examine the reasoning directly. All materials are public: Repository overview and methodology Exact prompt given to the AI AI’s original decision AI’s comparison after seeing the official decision Official final decision Official correction Public USPTO case search , where you can enter IPR2025-00030 The USPTO documents can be opened without creating an account. I would especially appreciate comments from patent lawyers, electrical engineers, and people who study AI evaluation. submitted by /u/hashiromer

Originally posted by u/hashiromer on r/ArtificialInteligence