Join the discussions about products powered by Cognite Data Fusion. Click the + CREATE TOPIC button in the menu bar to start the conversation.
Recently active
Hello, I have a SQL table that has 2000 rows total. The SQL DB Extractor 3.0 configured to extract all rows (Select *) and places them in a Raw table.The extractor log shows 2000 rows extracted.2024-08-02 00:00:05.558 UTC [INFO ] ThreadPoolExecutor-2_0 - No more rows for Caster_LabData_Narrow. 2000 rows extracted in 0.047 seconds2024-08-02 00:00:05.558 UTC [INFO ] ThreadPoolExecutor-2_0 - Reporting new success run: No more rows for Caster_LabData_Narrow. 2000 rows extracted in 0.047 seconds The Raw table UI shows it has only 1000 rows of data, not the 2000 I am expecting. How can I confirm the exact number of rows in the Raw table? I tried to select all rows from that raw table with a transformation and the preview only showed 1000 rows BUT I can’t tell whether that is a transform preview UI limitation or not.As a test, I tried to select all rows from a different raw table with 1993 rows with a transform and place the results in a 3rd raw table. The resulting 3rd table only shows
The workflow execution id is 1c3ef35b-c66e-4bdc-bf73-09480e291507.The workflow name is yggdrasil_domain_model_containers_transformations_including_raw_transformations. The tasks threw the following error right after the completion of first step, the error information is : 'reasonForIncompletion': 'Something went wrong. Please report this error to support@cognite.com and provide us with the workflow and task identifiers.', Project: Integral-develop
Hi,We have a large dataset containing more than 10,000 timeseries, and we are working on creating multiple subscriptions to accommodate them but again that can also be having 10k timeseries in sometime. Currently, we are using metadata with a specific property match as one of our filters, given that this metadata could also apply to 10,000 timeseries. we are considering adding a new filter based on createdTime. This would automatically include timeseries added after a specific time to a new subscription, avoiding the need to create new subscriptions manually.Here is the filter we attempted to use:"filter": { "and": [ { "equals": { "property": ["dataSetId"], "value": "12346576588" } }, { "range": { "property": ["createdTime"], "gte": 1719513000000 } } ]}We encountered an internal server error (request ID: b45316a2-07ac-9ec8-986f-75bd66d698e4) .Tried with creating createdTime as metadata as well, suspect there might be a limit
On behalf of Celanese, I’d like to know if there is any possibility on blocking download of single files on the CDF UI, if it is in a security category or a specified dataset. Thanks!
Hi,I see this error in UI when I open the data workflow in slb-forum project. In access-info, below permissions are listed:{ "workflowOrchestrationAcl": { "actions": [ "READ", "WRITE" ], "scope": { "all": {} } }, "projectScope": { "projects": [ "slb-forum" ] } } Can you please check if anything is missing?
While Loading Data to Power BI , it is taking huge time for loading transformed data.
Thus far, we have successfully loaded tables containing few million records from Delta Lake to CDF raw. However, we are encountering data skipping issues when we attempt to load more than ten million records from Delta Lake to CDF.I've just found the cause of all the records' failure to load.We are generating a sha2 key in our ETL process using the delta lake table primary keys. We don't have any problems with the sha2 key for small tables, but as we hit above 10 million records, we start to get key collisions, which causes some data to skip loading.In that instance, I made a new index field with incremental numbers like 1, 2, 3, and so on,. I then tried loading the data using the new index, and it now loads flawlessly. I was able to load 60 million records in CDF using this new method.
As per documentation, it is mentioned that 10,000 timeseries is the limit in each subscription. I see issue with count while adding timeseries using SDK.I have created timeseries subscription through SDK, it took more than 10000 timeseries. Is there any change in limit under subscription ?
Hello,I’ve been working on developing a Power BI dashboard that pulls the APM Data Model for InField to visualize our progress on checklists, however I’m running in to some road blocks. Currently for me to pull checklists, checklist items and measurements in power BI, I have to expand the related table columns within power BI rather than just being able to load each table individually and relate them to each other. Ideally I’d like to have the relationships to other tables as part of the model pull without needing to expand related tables as that’s creating significant folding in the data pull.Example:Loading the checklist table and expanding the checklistitems table and then the measurements table takes roughly 45 minutes to load due to the folding that occurs in the data load If I load these 3 tables separately, it takes under 1 minute - however I then don’t have any way to relate the values from each table to each other. The checklist items table does not have the checklist external
HiI am looking into using the is_new functionality in my transformations, but was wondering whether it will work in transformations where is use joins. If so, how would I write the query?In addition, I heard that you are working on releasing functionality to use is_new when reading from data models. When will this feature be ready?Thank you!Sebastian
Hello! Is there a way to push multiple timeseries in CDF timeseries within one single query in the DB extractor config.yml file? Or do we need multiple queries?Also, if my incremental field is a timestamp datetime format, is it working or not as I'm receving an error related to it. Regards,Raluca
Notify an instructor and ask them to give required consent to the application in the AAD (requires admin privileges): Share the link of your Grafana instance (https://<NAME OF DOMAIN>.grafana.net/) with the instructor The instructor will navigate to the link and click the box next to "Consent on behalf of your organization" and "Accept" Once the required permissions are granted sign in to https://<NAME OF DOMAIN>.grafana.net/ in a new session (e.g. incognito mode) and choose Sign in with MicrosoftPlease assign me required privileges.https://parthsinha.grafana.net/
Datapoints can have three main status codes: Good, Uncertain and Bad. https://developer.cognite.com/dev/concepts/reference/status_codesIs it possible to define a user defined status code? In our case, that would be Replaced. (I’m aware of the the sub category GoodEntryReplaced under Good, but it would be great to have Replaced as a main status code)
Good afternoon, I have to create a variable of type "query" in grafana, I have a data model called "DM_ALL_ASSET_QUALITY_PASSPORT"" with the "Asset_Quality_ListaPropiedades" view, which contains 3 columns that are: "Asset", which is the type of assets, "Header" which tells me if it is General or Metadata, and there is the "Property" column. which contains the properties that each type of assets belongs to, my problem is when I try to call the "Properties" depending on the assets, but I always get this type of error: "Parser: Syntax error: assets{metadata={Type='Manifold'}} | select {Property}" , I try to run the query "assets{metadata={Type='${Asset}'}} | select {Property}" I have tried in different ways but I have not been able to find someone who can help me, attached image. Thanks a lot @Aditya Kotiyal @HanishSharma @David Alvarez
Hello Team,We are trying to identify event type from data_modeling.data_models.list sdk method.as for sequences we get a response something like{ "container": { "space": "pgnig_space", "external_id": "SimulationResultSequences" }, "container_property_identifier": "data", "type": { "list": false, "type": "sequence" }, "nullable": true, "auto_increment": false, "name": "data"}So here we can identify sequnce by type.type= sequence for event type we get something like:{ "type": { "space": "slb-pdm-dm-governed", "external_id": "Event.entities" }, "source": { "space": "slb-pdm-dm-governed", "external_id": "Event", "version": "1_7", "type": "view" }, "name": "events", "direction": "inwards", "connection_type": "multi_edge_connection"} so here we don't get a type event. It refer s to a view and even if we go inside the view we cannot see any type specified as event.The only differ
We’re excited to invite you to Impact 2024, Cognite's first-ever User Conference, happening October 14-15, 2024 in Houston! This is your chance to connect with fellow professionals, innovators, and leaders from across industries as we explore the future of industrial AI and data innovation together.This two-day event is more than a conference – it’s a place for our community to learn, share, and grow. Whether you’re looking to tackle challenges, collaborate on innovative solutions, or just get inspired, Impact 2024 is the perfect platform to drive transformation and build meaningful connections. What’s in Store:Day 1: Industrial AI and the Digital JourneyTheme: Ambition Drives Outcomes, Data Foundation for ScaleDive into the world of Industrial AI and explore how to build a strong data foundation to drive real, scalable outcomes for your business. Day 2: Collaboration and Innovation with Cognite Data FusionTheme: Better Together: By Users, For UsersJoin deep-dive sessions with our prod
Hello Developers!In June, we launched the new Data Workflows service in Cognite Data Fusion. This service enables you to orchestrate Transformations, Functions, and more, bringing powerful capabilities to your data management processes.To help you get started, we've made some new content available:Cognite Learn: Check out our introductory video that covers the core concepts of data workflows. This is the first in a series of videos and guides we’ll be rolling out in the coming months. Jupyter Notebook in Fusion: Explore an example notebook in your CDF project, demonstrating how to leverage the features of the data workflows service. You can find it by navigating to Jupyter Notebook in the Data Management workspace in Fusion, and then go to the shared folder /Examples/data-workflows. We encourage you to try out Data Workflows and share your feedback with us either here or in the dedicated Hub group. We're continuously working to enhance the service with new features, so stay tuned for
We have a DM with many direct relations. A processing is directly liked to an asset with the name of the property being ‘equipment’. list shows the field as either just text or a dictionaryWe use the Cognite data modeling API to generate models and use pygen to run programs on it. I have noticed something strange while using pygen. Properties which are direct relations to other types sometimes are shown as text and other times are shown as NodeId objects.On close inspection, they appear as NodeId objectsIn the former case, while observing them in the Data Model view, the properties appear empty. Actually, they contain the external_id of the instance as a string which can be seen while querying with pygen. DM view showing seemingly empty cells Sometimes, instances spontaneously transition from one type to the ohter (NodeId → string or vice versa). In my pygen programs thus, I have to use the following function to keep the program running. from cognite.client.data_classes.data_modeling
In Cognite Data Fusion data models, instances (nodes and edges) are uniquely identified by their space and external ID.To simplify the user experience, Pygen includes a parameter called default_instance_space. You can set this parameter when generating a new SDK, allowing you to work with nodes and edges without specifying their space, as long as all nodes and edges share the same instance space.In earlier versions, if you didn’t specify the default_instance_space, Pygen would default to using the same space as the data model. Now, Pygen generates an SDK that requires users to specify the space when creating, retrieving, and deleting nodes and edges.The reason for this change is that having the schema (data model, views, containers) and data (nodes and edges) in the same space is considered an anti-pattern. Governance of a data model should be in a separate space, while data ingestion and consumption should occur in different spaces, typically one space per data source.
Hello Everyone! Exciting news! We're hosting Impact 2024, Cognite’s first User Conference, and we want YOUR use cases to shine. You can read more about Impact 2024 here. Why Share your Use Case?Your stories inspire. Share how your projects have driven real business impact, from boosting revenue to streamlining processes. We also welcome discussions on challenges you've faced and unresolved issues. This can spark meaningful conversations and potential solutions during the event.How to Get Involved:Submit your use case in this thread including it's impact. Based on the submissions, we will organize workshops to explore these use cases further. If your use case is selected, you will receive a free ticket to Impact 2024. We are committed to fostering a collaborative environment while respecting confidentiality and competition concerns, ensuring we all become better together. Mark Your Calendars:Submissions Open: April 25, 2024Submission Deadline: June 30, 2024Winners Announced: August 1, 2
Hi,From what I can tell from the documentation and what I can see from the datapoints we have in CDF the datapoints seem to be using milliseconds. My question is if CDF is currently able to handle microsecond timestamps or if this is functionality that would require additional development and whether or not this is on the current roadmap. Markus Pettersen
pygen version v0.99.30has just been released with a new approach to creating queries. This is experimental and we are looking for feedback. Check the documentation for more information.
Hello Everyone 😊, I'm now working on an endeavour that will integrate our organization's multiple old systems with Cognite Data Fusion (CDF). Although there are numerous advantages to CDF's current architecture, there are certain obstacles we must overcome in order to guarantee smooth data transfer between CDF & our more antiquated systems—especially those that weren't created with contemporary data platform in mind.I would appreciate your thoughts on the following specific queries and worries:Data Connectivity: How can CDF and older technologies that use antiquated communication formats or protocols create and sustain reliable connections with one other?🤔 Have anyone of you run across problems with particular kinds of systems?🤔 If so, tell us about it.Data Consistency and Quality: When importing data into CDF from older systems, how can you maintain and guarantee data consistency?🤔 Exist any particular Cognite ecosystem tools or procedures that support preserving data accuracy
Hi There,We would like to confirm if the indexing mechanism in the CDF operates in the same manner as it does in a relational database. Specifically, we need to understand the trade-offs of using indexes carefully in CDF. In relational databases, indexes occupy space on disk and memory when in use, which can be problematic if space or memory is limited. Additionally, maintaining indexes during data insertions, updates, or deletions can slow down these operations and lock tables (or parts of tables), potentially affecting query performance.Given these considerations, do we need to manage indexes in CDF with the same level of caution as in relational databases?Disadvantages of having an index in a relational database:Space: Requires additional disk/memory space. Write speed: Slows down INSERT, UPDATE, and DELETE operations.
Hi:In a cognite video I saw that they talked about image and video contextualization, where CDF can recognize and identify patterns in images, is this possible?.I would be very interested in offering this service to my clients.