AI Training and Copyright: Why Your Business Should Care
You may not be building an AI model, but your business could still be affected by the growing dispute over how AI systems learn from copyrighted material. If your team uses ChatGPT, Claude, Gemini, image generators, transcription tools, or AI features inside business software, the legal and operational questions are no longer limited to large technology companies.
The immediate issue is whether AI companies can train large language models using books, news articles, images, and other copyrighted works without permission. A recent United States government brief supported OpenAI’s position in a lawsuit brought by The New York Times, arguing that restricting AI development through an overly narrow interpretation of fair use could slow technological and scientific progress. However, the brief is not a court ruling, and it does not settle the question for businesses in Malaysia.
For you, the practical question is simpler: how can you use AI productively while reducing the risk of submitting confidential material, producing copied content, or relying on unclear ownership rights?
TL;DR
AI copyright rules are still developing, and the United States government’s brief supporting OpenAI is not a final legal decision. Your safest approach is to create clear internal rules, protect confidential information, verify AI-generated work, and keep records of what your team uses and publishes.
What This Means
Large language models are trained on very large collections of text and other content. These collections can include published articles, books, websites, code, images, and media. The dispute is whether using copyrighted material during training is legally allowed without permission.
One argument is that training is transformative: the AI system does not simply display an entire book or article to every user. Instead, it analyses patterns and uses them to generate new responses. This is similar to the argument that a person may read many books before writing an original explanation.
The opposing argument is that AI companies may be commercially benefiting from copyrighted works without paying or obtaining permission. Copyright owners also worry that AI systems can reproduce distinctive passages, imitate styles, or compete directly with the creators whose work helped build the systems.
The TechCrunch report describes a 20-page brief from the US administration supporting OpenAI’s position and referring to the importance of maintaining American leadership in artificial intelligence. The same report notes that the brief has no direct authority over the US district court hearing The New York Times lawsuit. Read the source report.
There is also an important distinction between training an AI model and using an AI tool. Your business is unlikely to train a general-purpose model from scratch. More likely, you will upload a sales document, ask for a product description, summarise a customer message, generate an image, or use an AI assistant within an existing application. Each activity creates different questions about confidentiality, ownership, accuracy, and permitted use.
AI can help you create faster, but speed does not remove your responsibility to check what goes into the system or what comes out of it.
How This Applies to Malaysian SMEs
Consider a Malaysian retailer that asks an AI tool to write product descriptions. If the employee copies text from a competitor’s website, a supplier catalogue, or a magazine article into the prompt, the business may be creating a separate copyright and confidentiality problem. The safer process is to provide your own product facts, dimensions, warranty details, tone, and target audience, then ask the tool to draft fresh wording. A staff member should compare the result with the source material before publishing.
A service business may use AI to prepare proposals, quotations, customer replies, or social media posts. This can be useful when you have a busy sales team, but the prompt might contain customer names, telephone numbers, project details, architectural drawings, medical information, or internal operating procedures. Before using an external AI tool, check whether the provider retains prompts for service improvement or allows business data controls. If you cannot confirm the treatment of sensitive information, remove identifying details or use an approved business environment.
Manufacturers and distributors face another practical concern. Your team may upload supplier manuals, technical drawings, product photographs, or imported documentation to create training notes. Some of these materials may be licensed only for internal use. AI-generated summaries do not automatically change those restrictions. Keep the original licence terms, identify the approved audience, and do not assume that a generated document can be shared publicly just because the wording is new.
Marketing agencies, tuition centres, restaurants, clinics, property firms, and professional practices also need to review image generation. An AI-created image may resemble a real person, brand, character, artist’s style, or protected design. You should not use a generated visual in a campaign simply because the tool produced it. Ask whether the tool provides commercial-use terms, retain the generation record, and arrange human review before publication.
For Malaysian SMEs, the legal position may involve more than copyright. Personal data handling, customer confidentiality, contract terms, trade secrets, and sector-specific obligations may also apply. The Personal Data Protection Department provides information on Malaysia’s personal data protection framework. If your business handles health, financial, education, or employment information, obtain professional advice before introducing AI into those workflows.
A Simple Risk Map for Your Team
| AI activity | Main concern | Basic control |
|---|---|---|
| Drafting marketing copy | Copied wording or unsupported claims | Use your own facts and conduct a human review |
| Summarising customer documents | Confidential or personal information | Remove identifiers and use an approved tool |
| Creating images | Similarity to protected people, brands, or artwork | Check usage terms and approve visuals before publishing |
| Writing software code | Unclear licence obligations or insecure code | Review licences, test code, and scan for vulnerabilities |
| Preparing business decisions | Errors, bias, or fabricated information | Require a responsible person to verify the output |
Practical Takeaways
- Create a one-page AI policy. State which tools employees may use, what information they must not upload, and who approves public content.
- Classify business information. Mark documents as public, internal, confidential, or restricted. Only public and approved internal material should normally enter general-purpose tools.
- Use original inputs. Give AI your product facts, process notes, and approved brand guidelines rather than copying large sections from third-party content.
- Keep a review step. Check facts, calculations, translations, legal statements, product claims, and tone before anything reaches a customer.
- Keep records. Note the tool used, date, prompt purpose, source materials, and reviewer for important proposals, campaigns, and operational documents.
- Read provider terms. Pay attention to training settings, retention, business data protection, ownership language, and permitted commercial use.
- Ask suppliers and agencies questions. Your contracts should clarify who is responsible for licences, generated assets, confidential information, and correction of errors.
- Escalate sensitive cases. Speak with a qualified lawyer or privacy adviser when content involves major campaigns, customer data, employee records, licensed material, or potential disputes.
The Bigger Picture
The dispute involving OpenAI and The New York Times is part of a wider adjustment in how society defines authorship, training, originality, and permission. The US government’s intervention may influence the policy conversation, but it does not remove uncertainty for companies operating elsewhere. Courts, regulators, publishers, software developers, and customers will continue shaping the rules.
For your business, waiting for every legal question to be answered is unnecessary. You can already build responsible habits that remain useful under several possible outcomes. If future rules require clearer permission for training data, you will be better prepared if your own content sources and supplier terms are documented. If customers demand transparency, you will be able to explain how AI was used. If a provider changes its policy, you will know which workflows need review.
The long-term advantage will not come from using AI everywhere. It will come from using it in the right tasks with clear boundaries. Routine drafting, classification, internal search, translation, meeting summaries, and first-pass analysis may benefit from AI. Final judgement, customer trust, compliance, and accountability should remain with people you have assigned to those responsibilities.
Start with one controlled workflow. Choose a low-risk task, write down the approved information, require a review, and monitor the result for a month. Then improve the process before expanding it. That approach gives you practical value without treating uncertain copyright questions as a reason to ignore basic business controls.
Ready to Streamline Your Operations?
Your business should run itself. AutoRunBiz deploys AI agents to automate your daily operations — WhatsApp orders, invoicing, customer follow-ups, and accounting. Book a free 15-min ops audit to see where automation fits your business →
