Ask ChatGPT this exact question and it will give you a reasonable-sounding yes, with a few caveats about double-checking your work. It won't cite Treasury Circular 230. It won't mention that the IRS Office of Professional Responsibility issued guidance on this specific question in June 2026.
It almost certainly won't tell you that the same tool answering you hallucinates on more than half of specific legal citations it's been tested on. That's the whole problem in one search box: the question keeps getting asked inside the tool the question is about, mid-doubt, usually right before something goes into a client file.
Here's the honest line. Yes, use it to explain a concept, draft a plain-English summary, or tighten a memo you already researched. No, don't let it generate a citation you haven't independently checked, and no, don't paste client data into a version that trains on what you type.
Everything below is where that line actually sits, and why.
A general-purpose model is good at exactly the tasks that don't require it to be a primary source. Ask it to explain the difference between a step transaction and a substance-over-form challenge and it will usually give you a clear, accurate refresher, the kind of thing a mid-level associate half-remembers from a CPE course two years ago. That's a legitimate use, and it's the one most CPAs are already comfortable with.
The same is true for translation work. Once you've done the research and reached a position, a general model is genuinely useful for turning a technical analysis into the plain-English version a client can act on, or for tightening a memo's prose without touching the substance.
It's also decent at a first-pass issues list: paste in a fact pattern and ask what tax questions it raises, not what the answers are, and you get a reasonable checklist to start real research from. Document summarization works the same way. Paste in a lease, an operating agreement, or a closing statement and ask for a plain summary of the terms, and the model is reading text you supplied rather than recalling a legal authority from memory, which is a very different reliability profile.
Notice what all of these have in common: the model isn't the source of the facts or the law. You are, or the document you pasted in is. That's the boundary.
The moment a general model has to produce a citation from memory rather than work with something you gave it, the reliability profile changes completely. This isn't a vague warning. It's a specific, well-documented failure pattern with three parts.
• Fabricated authority: Ask a narrow enough question and a general model will sometimes return a Revenue Ruling number that doesn't exist, a Tax Court case with a plausible name and a docket number that goes nowhere, or a Code section reference that's confidently wrong by one subsection. It doesn't read as a guess. It's formatted exactly like a real citation, which is what makes it dangerous: nothing about the output signals uncertainty.
• Confident errors on numbers that change every year: A model's knowledge has a cutoff, and tax thresholds, phase-outs, and bonus depreciation percentages move annually, sometimes by statute mid-year. A model asked about bonus depreciation percentages in 2026, for instance, might confidently describe the old TCJA phase-down schedule (40% declining to 20% to 0%) without knowing that the One Big Beautiful Bill Act restored 100% bonus depreciation for property acquired and placed in service after January 19, 2025. That's not a hallucinated citation. It's a real rule, stated with total confidence, that stopped being current law over a year ago.
• Thin state coverage, and no way to check currency: Stanford's hallucination research found error rates climb specifically on lower-court and less-prominent authority, which describes most state tax questions almost exactly: state DOR rulings, state tax court decisions, and older guidance that isn't famous enough to be well-represented in training data. And a general chatbot has no citator function. Westlaw's KeyCite and Lexis's Shepard's exist specifically to tell you whether a case has been overruled or a ruling superseded; ChatGPT has no equivalent, and it won't volunteer that a rule it just stated has since changed.
This is the exact line a tool built for tax research has to hold, and it's the specific gap Bizora is built to close: primary-source coverage across all 50 states, not just federal, so the state and lower-court authority general models handle worst is exactly where a purpose-built tool is strongest.
None of this is a new legal problem. It's an old rule meeting a new tool. Treasury Circular 230 (31 CFR Part 10) already required the things that AI use tests.
Section 10.22 requires due diligence in determining the correctness of anything represented to the IRS or a client. Section 10.35 requires the competence to actually recognize when an answer is wrong.
Section 10.37 is the one that matters most here: written advice can't rest on unreasonable assumptions or on representations where reliance would be unreasonable. There's no carve-out in any of these sections for "the tool generated it." A citation is either right or it isn't, and the practitioner who signs the memo owns that answer regardless of where it came from.
The IRS made this explicit in June 2026, when the Office of Professional Responsibility issued Alert 2026-19, its first substantive AI guidance. The language leaves no room for a different reading: "due diligence requires verifying the accuracy of facts, citations, and calculations produced by AI," and "practitioners cannot rely solely on AI."
The alert also points to what happens when nobody does: in 2025, Deloitte Australia delivered a government report with invented quotes and non-existent references, and ended up partially refunding its fee once the errors surfaced. There's no "the AI said so" defense in any of this. There never was one for a junior associate's bad citation, and there isn't one here.
That principle has already reached courtrooms. In Mata v. Avianca (S.D.N.Y. 2023), attorneys were sanctioned $5,000 for a brief containing six ChatGPT-fabricated cases with invented quotes attributed to real judges. Their actual mistake wasn't using the tool. It was never checking the output, then defending it after being told.
Two more incidents landed specifically in tax venues: the U.S. Tax Court struck a filing in Thomas v. Commissioner after three of four cited cases turned out not to exist, and a Minnesota Tax Court case, Delano Crossing 2016, LLC v. County of Wright, ended with a government attorney referred to the state's professional responsibility board for citing cases that were never real.
No CPA has been formally disciplined for this yet in any public record. That's a timing gap, not a reason to assume the rule doesn't apply.
Everything above assumes the only risk is a wrong answer coming back. There's a sharper problem that happens before the model ever responds: what happens to the client data typed into the prompt to get there.
Free and Plus-tier ChatGPT train on conversation inputs by default, unless someone manually opts out in data controls. Most CPAs pasting a client's numbers into a chat window have never made that election, and many don't know it exists. Business, Enterprise, and API-based access don't train on customer data by default and support zero data retention, which is a meaningfully different product for this purpose even though the interface looks the same.
This isn't a soft concern. IRC Section 7216 makes it a criminal misdemeanor for a preparer to knowingly or recklessly disclose or use tax return information for an unauthorized purpose, and Section 6713 imposes a parallel civil penalty for the same conduct even without that intent standard.
Treasury's regulations define "disclosure" broadly enough, "in any manner whatever," that pasting a client's return details into a consumer chatbot is a plausible violation, not a stretch reading. The IRS hasn't issued a definitive ruling on this exact fact pattern, but the practitioner consensus is to treat it as a real, live exposure rather than a technicality to wave off. If your firm has a client using AI note-takers, transcription tools, or chat assistants that touch return information, this same question applies there too.
None of this adds up to "don't use ChatGPT." It adds up to a short list of habits that make the difference between using it well and creating a problem for yourself.
Never paste, on a consumer tier: Social Security or EIN numbers, account numbers, a client's full name combined with return details, or any combination specific enough that someone could reconstruct whose return it is. If the work requires that level of detail, use a Business, Enterprise, or API tier with a no-training commitment, or don't use AI for that step at all.
Verify before it goes anywhere client-facing: Treat every citation, Code section, case name, or dollar threshold the tool gives you as unconfirmed until you've pulled it up yourself on a primary source, IRS.gov, the actual regulation, the court's own opinion, not a summary of one. If you can't locate it independently in a couple of minutes, assume it's fabricated rather than assuming it's probably fine.
Escalate to a primary-source tool when it matters. General research, drafting, and issue-spotting are fine in a general chatbot. Anything that touches a threshold or percentage that changes annually, anything state-specific, anything that will appear in a signed memo or a position taken on a return, and any multi-step reasoning across more than one authority is exactly where the hallucination rate climbs and where a tool built to cite primary sources earns its cost.
The honest pitch here isn't "stop using ChatGPT." It's that the two failure modes above, fabricated citations and unverified disclosure, are exactly what a tool built for tax research is supposed to close, not something a general chatbot can patch with a better prompt.
Bizora traces every answer to the specific IRC section, Treasury regulation, or ruling behind it, with the reasoning visible through View Steps, so the citation is something you click and check rather than something you have to go find yourself. If you're weighing what actually matters when comparing tools on this front, the guide to choosing an AI tax research assistant covers primary-source access and reasoning transparency as selection criteria rather than marketing language. And once an answer is verified, the guide to structuring a tax research memo covers documenting it properly, since a well-documented memo is what Section 10.22 due diligence looks like on paper when someone asks to see it.
Ask Bizora the question that opened this article, and the answer isn't different because it's more confident. It's different because it comes with a citation attached that you can actually verify.
No. There's no rule against the tool itself, and plenty of tax research tasks are genuinely fine to run through it. The obligation is on the use: independently verify any AI-generated citation, calculation, or conclusion before it reaches client work, and keep client data out of tools that train on what you type.
Treasury Department Circular No. 230, codified at 31 CFR Part 10, sets the standards of practice for attorneys, CPAs, enrolled agents, and other practitioners representing taxpayers before the IRS. It covers due diligence, competence, and the requirements for written tax advice, and it's the primary source of professional discipline for the profession.
Stanford RegLab research found general-purpose models hallucinate on 58% to 88% of specific legal queries, worse on lower-court and less-prominent cases, which describes most state tax authority. A separate study put GPT-4 specifically around 43%, against 17% to 33% for tools built for legal research with retrieval systems.
Yes, for the tasks that don't depend on it being a primary source: explaining a concept, tightening a memo, summarizing a document you pasted in yourself, or generating a first-pass issues list. The confidentiality risk disappears without client data involved, but the citation-accuracy risk doesn't, so anything the tool asserts as law still needs independent verification.
Attorneys in Mata v. Avianca were sanctioned $5,000 for a brief with six fabricated cases. Since then, similar incidents have reached tax forums specifically, including a U.S. Tax Court filing and a Minnesota Tax Court case where the attorney was referred to a professional responsibility board.
It's a genuine, largely unexamined risk. Free and Plus-tier ChatGPT train on inputs by default, and IRC Section 7216 makes unauthorized disclosure of tax return information a criminal matter, with Section 6713 imposing a parallel civil penalty. Business, Enterprise, and API-based access typically don't train on inputs by default, which meaningfully lowers, but doesn't erase, the exposure.
Bizora AI turns hours of manual research into seconds, with every answer backed by primary source citations. Start your 7-day free trial. No credit card required.
Start Free Trial