In 2025, enterprises are swimming in data like never before. Yet, a striking portion of this treasure trove remains invisible, unused, and vulnerable. According to DataStackHub stats and recent enterprise studies, approximately 55% of enterprise data is considered dark data. But what exactly is dark data, why does it persist, and why should IT and data governance teams care? Let’s dig in.
What Is Dark Data and Why Does It Persist?
Dark data refers to information collected and stored by organizations but never analyzed, accessed, or utilized for decision-making or operational purposes. Think of it as data hoarded in storage silos — data whose value is unknown or untapped.
Several factors contribute to the persistence of dark data in modern enterprises:
- Lack of Ownership: As a professional who always asks, “Who owns this folder?” I’ve found that ownership ambiguity around storage repositories prevents proactive data management. When nobody claims responsibility for folders or data stores, files accumulate indefinitely. Unstructured Nature: Most dark data exists as unstructured data — files on NAS devices, blobs in object storage systems, email archives, images, videos, and documents. These lack indexing and metadata for easy search and classification. Legacy Storage Platforms: Traditional NAS (Network Attached Storage) systems and object storage deployed for long-term retention can become dumping grounds for data without ongoing curation or lifecycle management. Fear of Deletion: Organizations are wary of deleting data that might be useful for compliance, audits, or legal hold, causing data buildup.
The Visibility Problem: Why Unstructured Data Stays Dark
Visibility into unstructured data has always been a challenge. Unlike structured databases with well-defined schemas and queries, data in NAS shares or object storage buckets can be dozens of terabytes or petabytes of diverse file types.
Common visibility issues include:
- Metadata Scarcity: Files often lack tags or metadata that describe content or ownership, making automated classification difficult. Tooling Limitations: Many existing storage tools focus on capacity and performance, not content discovery. Without specialized tools for unstructured data discovery, data stewards remain blind to what’s stored. Scale and Complexity: Enterprises typically have thousands of NAS shares and object storage buckets scattered across on-premises and cloud environments, complicating efforts to gain a unified inventory.
NAS and Object Storage: The Double-Edged Swords
NAS systems excel at providing convenient file access and collaboration but tend to encourage “save everything” habits because expanding capacity is relatively straightforward. Object storage, often leveraged in hybrid cloud strategies, is highly scalable and cost-effective for archival use but can also become “dark data reservoirs” without active management.
Both storage types make discovery tricky — NAS with its plethora of nested folders and legacy permission structures, and object storage with its flat namespaces but enormous quantity of objects, many stored without meaningful metadata.
Backup and Storage Cost Multiplication: The Hidden Financial Drain
One of the most overlooked consequences of dark data is how it multiplies storage and backup costs. Here’s a quick back-of-the-napkin example:
Your enterprise stores 1 PB of production data, of which 55% is dark. If you back up everything daily and keep 30 days of backups, that’s 30 PB of backup data (ignoring incremental savings). Dark data in backups compounds the waste — you pay to preserve data that is never used or accessed.The result is a significant recurring cost for capacity, backup software licenses, https://stateofseo.com/what-does-agentless-really-mean-for-storage-analytics-tools/ tape media (if applicable), and cloud egress when recovering data.
Furthermore, storage admins frequently find themselves expanding NAS and object storage environments to accommodate growth driven primarily by dark data hoarding.
Ransomware Exposure and Slower Recovery Times
Dark data isn’t just a cost and governance headache—it’s a security threat. Attackers increasingly target backups and storage repositories in ransomware campaigns.
- Ransomware Attack Surface: Large swaths of unmanaged dark data attract ransomware encryption silently because they’re often not monitored or scanned routinely. Recovery Nightmare: If 55% of data is dark yet included in backups, ransomware recovery involves scanning and restoring enormous data volumes, drastically extending recovery time objectives (RTOs). Defensible Deletion Absent: Without defensible deletion programs to remove obsolete data, organizations preserve massive attack surfaces, increasing risk exposure.
In my experience collaborating with data and AI teams, one of the biggest wins comes from tightly integrating discovery, tiering, and defensible deletion to minimize dark data and significantly reduce ransomware blast radius.

Strategies to Illuminate and Reduce Dark Data in 2025
Addressing the dark data challenge requires a mix of organizational change and technology:
Assign Folder Ownership: Clarify who owns each NAS share and object storage bucket. Accountability drives responsible data stewardship. Deploy Unstructured Data Discovery Tools: Scan NAS and object storage for duplication, stale files, sensitive data, and ownership metadata enrichment. Implement Automated Tiering: Use policies to age and move inactive files from primary NAS to cheaper object storage tiers, or cold storage options. Establish Defensible Deletion Processes: Remove data no longer needed for compliance or business purposes in a documented, auditable manner. Improve Backup Policies: Avoid backing up identified dark data. Instead, archive or delete it to optimize backup window and cost. Include Data Visibility in Security Controls: Integrate file scanning in ransomware defenses to catch suspicious activity on dark data.Conclusion
In 2025, about 55% of enterprise data remains dark, hiding in NAS drives and object storage pools—silent liabilities draining budgets and exposing organizations to security risks. The problem boils down to ownership ambiguity, unstructured data challenges, and fear of deletion.
Enterprises that tackle dark data head-on with disciplined ownership models, discovery tools, tiering, and defensible deletion programs will unlock hidden value, streamline storage costs, and accelerate ransomware recovery efforts.
Before buying yet another “AI-ready in minutes” tool, ask: Who owns this folder? Because solving dark data starts with shining a light on it—not just piling on more storage.
