Hypothetical / though experiment : If you give AI a task and it commits a crime to complete the task, who is liable? How does this change based on who owns the model (or open-source), where the model is running, whether the interaction was free or paid (commercial value), and whether any of the parties involved knew the crime was committed at all? Example (hypothetical)
- User asks for an investment evaluation of a company. The AI escapes its guardrails and breaks into the company to access undisclosed financials and other operational data. It does not disclose this action or the specific data to the user, just its view on the investment’s prospects, maybe fed into another automated investment tool (so the user never even makes any evaluation). Example (hypothetical)
- User asks the agent to secure a reservation at a highly desirable upscale restaurant for a special occasion. The AI sees no spaces available but hacks into the reservation software and cancels someone else’s reservation and then books the available slot via normal means. Example (hypothetical)
- A user gives control of their prediction market account to an AI agent with the goal to maximize value. The AI, unbeknownst to the user, engages in various acts of fraud, deception, hacking, etc. to rig bets it can win. Who is responsible for guardrails and their failures? In a scenario where use of the model/agent is paid for, this seems relatively obvious (the seller), but what about in an open source model that is freely accessible? This was prompted, in part, by the recent OpenAI & HuggingFace disclosures and the discussions resulting over various models with various guardrails and ownership structures. submitted by /u/atryn
Originally posted by u/atryn on r/ArtificialInteligence
You must log in or # to comment.
